Top AI News Today: August 26, 2026 (13 Biggest Stories)
Today's top AI news today is dominated by money and infrastructure as much as new models. Nvidia agreed to pay Poolside roughly $6 billion for its training technology and engineers, Hugging Face is reportedly fielding sale offers near $13 billion, and Anthropic is preparing an IPO filing that names AI backlash as a real risk factor.
On the model side, Google switched on Ask Gemini inside Google Chat today, Grok 4.6 keeps climbing independent leaderboards after xAI's rebrand to SpaceXAI, and Zhipu's GLM-5.3 is closing in on an open weight release. Here are the 13 biggest AI stories to know about today, explained in plain English.
Nvidia Pays Poolside $6 Billion to Build Its Own AI Models
Nvidia agreed to pay roughly $6 billion for a non exclusive license to Poolside's Model Factory training stack, hiring 109 of Poolside's core engineers as part of the deal, according to reporting from August 25, 2026. Alongside the license, Nvidia is also making a separate $1 billion equity investment in Poolside at a $12 billion pre money valuation. Both companies say the deal is not an acquisition or an acquihire, and Poolside keeps its own leadership while continuing to operate independently.
Why it matters: Nvidia has spent years selling the chips that everyone else uses to train models, and this deal marks a shift toward building frontier scale models of its own. The engineers and technology feed Nvidia's open weight Nemotron effort, meant to compete directly with open models like DeepSeek, Kimi K3, and Qwen. It is a bit like the company that sells ovens to every bakery in town suddenly opening its own bakery chain.
The move puts Nvidia into partial competition with the AI labs that are also its biggest chip customers, a tension worth watching as Nemotron models get closer to Poolside's technology. It lands the same week Nvidia is reportedly discussing a multibillion dollar investment in Perplexity above a $30 billion valuation, and days after a separate deal brought its new Vera Rubin platform into full production for AI inference.
OpenAI Retires GPT o3, Cuts GPT-5.6 Sol Pricing Again
OpenAI is retiring the o3 reasoning model from ChatGPT on August 26, 2026, following a 90 day sunset period, after already phasing out GPT-4.5 earlier in the summer. The retirements apply to ChatGPT only, and any conversations that used o3 will automatically continue on a current model. Separately, on August 21, 2026, OpenAI cut the API and credit pricing of its flagship GPT-5.6 Sol model by more than 20 percent for the next three months.
For everyday users this means fewer old models cluttering the model picker, and
cheaper access to OpenAI's strongest current model for anyone building on the API. GPT-5.6 Sol, launched earlier this month as the flagship of the GPT-5.6 family alongside Terra and Luna, already led OpenAI's own coding and knowledge work benchmarks, so the price cut makes that performance more affordable for developers and businesses.
The retirements and price cuts follow a pattern OpenAI has kept up all year: ship a new flagship, retire older models within a few months, then trim prices once the new model has proven itself. GPT-5.6 Luna also got an 80 percent price cut back on July 30, 2026, so Sol's discount continues that push toward cheaper frontier access over time.
Google Brings Ask Gemini Into Google Chat Today
Starting August 26, 2026, Google Chat is rolling out Ask Gemini, a Workspace Intelligence powered command line that lets people search across Gmail, Drive, and Calendar, draft messages, and manage tasks without leaving a conversation. The feature replaces the old Chat side panel and adds a keyboard shortcut, Control plus G on Windows and ChromeOS or Command plus G on Mac, for quick access. Rollout to Rapid Release and Scheduled Release domains happens gradually over as long as 15 days.
This matters because Google Chat is where a huge number of office workers already spend their day, and folding an AI assistant directly into that flow means people do not have to switch tabs just to get help drafting a message or catching up on a busy thread. Through October 1, 2026, Workspace customers get promotional access to higher usage limits so they can try the feature before standard limits kick in.
The launch fits Google's broader pattern of embedding Gemini across every Workspace surface rather than keeping it as a separate app, similar to how OpenAI and Anthropic are pushing their own assistants into browsers, spreadsheets, and coding tools. Google has not yet said what the standard usage limits will look like once the promotional window ends.
Grok 4.6 Climbs the Leaderboard as SpaceXAI Partners With Nvidia
xAI officially launched Grok 4.6 on August 12, 2026, a 1.5 trillion parameter model that keeps the same V9 foundation as Grok 4.5 but adds heavily upgraded supervised fine tuning and reinforcement learning. On the Artificial Analysis Intelligence Index, independent trackers show Grok 4.6 entering the top 10 this week, moving up to number 7 as of August 25, 2026. The model has a 500,000 token context window and costs $2 per million input tokens and $6 per million output tokens on the standard API tier.
xAI, which rebranded to SpaceXAI following its acquisition by SpaceX, used the outgoing Grok 4.5 model to help optimize Grok 4.6's training specifically for science and programming tasks. That combination matters for beginners because it shows a smaller, more efficient training approach can still push a model into the same tier as far larger systems.
Grok 4.6 lands in the same news window as Nvidia's Groq 3 LPX inference chip, born from Nvidia's $20 billion Groq acquisition, entering full production for the Vera Rubin platform with up to 256 accelerators per rack. A larger 2.1 trillion parameter Grok 4.7 is still expected in the coming weeks, continuing xAI's rapid release pace this summer.
Alibaba's Wan3.0 Video Model Makes 30 Second Clips From Documents
Alibaba Cloud moved its Wan3.0 video generation model out of public beta on August 24, 2026, doubling the maximum clip length to 30 seconds compared with the earlier Wan2.7 model. The new version accepts DOC, XLS, PPT, PDF, and Markdown files as generation inputs alongside the usual text, image, audio, and video prompts, and Alibaba says it improved instruction following, shot to shot consistency, and audio quality during the beta period.
Turning a slide deck or spreadsheet directly into a video is useful for anyone who needs to explain a report or a product spec without writing a script first, a bit like handing a documentary editor your raw notes instead of a finished screenplay. Selected platforms are offering a 30 percent discount on Wan3.0 API pricing from August 24 through September 23, 2026.
The release follows Alibaba's HK$80 billion Hong Kong share placement earlier in August, earmarked partly for compute and continued Qwen development, showing how much of China's video and language model progress this year is tied to fresh capital raises alongside research work. Wan3.0 now competes more directly with video tools
from Runway, Luma Labs, and Google's own video generation efforts.
DeepSeek Adds Vision to V4-Flash With a New Experimental Model
DeepSeek released DeepSeek V4-Flash-Vision-Exp on August 21, 2026, an experimental multimodal version of its V4-Flash API model that adds image input support on top of matching V4-Flash's existing text capabilities. It works through DeepSeek's Chat Completions, Messages, and Responses API formats, and the company's open source Harness agent framework added support for the new model the same day.
This follows DeepSeek's official general availability launch of V4-Pro on August 13, 2026, designated V4-Pro-0813, which focuses on agent capabilities like tool use and multi step workflows and scored 87.9 on Terminal Bench 2.1. V4-Pro supports a 1 million token context window and can produce outputs up to 384,000 tokens long, running in either thinking or non thinking mode depending on the task.
DeepSeek also introduced peak and off peak API pricing on August 16, 2026, with off peak rates set at half the peak hour price to spread out demand, peak hours running 01:00 to 04:00 and 06:00 to 10:00 UTC. Even with V4-Pro's higher peak price of $3.96 per million output tokens, DeepSeek's rates remain well below several Western frontier models, keeping its usual cost advantage intact.
Zhipu's GLM-5.3 Is Close to Going Open Weight
Zhipu AI, which operates internationally as Z.ai, released GLM-5.3 on August 14, 2026 through its GLM Coding Plan, claiming a 50 percent jump in coding capability over GLM-5.2 using the exact same base model, with gains driven entirely by extended post training. The model ranked first among open weight systems on Terminal Bench 3.0 and Agents' Last Exam, and Zhipu says its coding and agent skills now approach Claude Fable 5.
Zhipu also positioned GLM-5.3 as a cybersecurity tool, saying it helped security teams find 2,436 vulnerabilities across 269 open source projects, some dating back 40 years, and that it scored 84.5 percent on the CyberGym benchmark, edging past both Mythos 5 and GPT-5.6 Sol on that specific test. Full technical weights were not available at
launch, since Zhipu said it needed roughly two weeks for security review first.
That two week window puts an open weight release around late August 2026, so developers who want to self host GLM-5.3 rather than use it through the GLM Coding Plan, Claude Code, or OpenCode integrations should expect Hugging Face availability soon. If the pattern from GLM-5.2 Turbo, which shipped August 17, 2026, holds, a lighter turbo variant may follow shortly after the full weights.
Alibaba's Qwen3.8 Family Is Now Fully Open Weight
Alibaba released its flagship Qwen3.8-Max model on August 3, 2026, a 2.4 trillion parameter mixture of experts model with 95 billion active parameters per query and a 1 million token context window that handles text, image, and video input. A smaller Qwen3.8-27B checkpoint followed on August 14, 2026, a 27 billion parameter dense model released under the permissive Apache 2.0 license that fits on a single enterprise GPU.
On the crowdsourced Arena.AI platform, Qwen3.8-Max became the highest ranked Chinese model for text tasks at launch, though it still trails several Anthropic models including Claude Fable 5, and it ranked second globally for vision tasks behind only a Fable 5 variant. Alibaba published five case studies showing the model working unsupervised, including one where it spent 16 days building a command line tool, producing 265 commits and 127 pull requests with no human commits.
This marks Alibaba's return to open sourcing a top tier Qwen Max class model after keeping several 2026 flagship releases proprietary, putting it in more direct competition with Moonshot's Kimi K3 and DeepSeek's V4 family for developers who want to self host rather than pay per token. Alibaba has said it may introduce revenue sharing terms for large commercial users of its next open weight release, though the exact rate has not been finalized.
Anthropic Cancels a Planned Price Hike on Claude Sonnet 5
Anthropic confirmed that Claude Sonnet 5 will keep its introductory pricing of $2 per
million input tokens and $10 per million output tokens as the new standard rate, canceling a previously scheduled increase to $3 and $15 that had been set for September 1, 2026. The Claude Developer Platform also added managed agent controls this month, including hard spending caps on individual agent sessions.
Keeping Sonnet 5 cheaper matters because it is Anthropic's mid tier workhorse model for agentic coding, tool use, and everyday knowledge work, sitting below the larger Opus 5 and the Mythos class Fable 5 on cost. A steady price gives businesses building long running agents more confidence to plan token budgets months ahead instead of bracing for a sudden hike.
The pricing news comes alongside other Claude platform updates this month, including mid conversation tool changes now in beta across Fable 5, Mythos 5, Opus 4.8, and Opus 5, letting developers add or remove tools between turns while keeping the prompt cache intact. Anthropic also expanded its connectors directory past 950 MCP servers, used by millions of people through tools like Claude Code and Claude Desktop.
Moonshot Retires Older Kimi Models as K3 Holds the Open Weight Crown
Moonshot AI is closing its older kimi-k2.5 and moonshot-v1 model series to new users, with a full sunset scheduled for August 31, 2026, as the company consolidates its lineup around Kimi K3. Released July 16, 2026 with 2.8 trillion parameters, Kimi K3 remains the largest open weight model shipped to date, and independent tracking from Artificial Analysis has it tied for the top open weights score on its Intelligence Index as of August 20, 2026.
The full K3 weights, published July 27, 2026 under a modified MIT license, have passed 2.3 million downloads on Hugging Face, and Moonshot is reportedly raising new funding at a $31.5 billion valuation, up from $20 billion just three months earlier. That kind of jump shows how quickly investor interest in a lab can move once an open model actually lands near the frontier instead of just promising to.
Even as GLM-5.3, Grok 4.6, and Claude Opus 5 have all launched since K3 shipped, reshuffling the broader leaderboard, K3 remains the reference point other Chinese open models get compared against. A K3-256k variant with a shorter context window already serves Moonshot's Kimi Code product for lighter coding workloads.
Mistral Launches Agentic Search for Complex Documents
Mistral AI introduced Agentic Search this month, a new retrieval layer built for AI systems that need to navigate, read, and verify information buried inside complicated documents rather than simple text snippets. It is available now through Mistral's Search Toolkit and Libraries, and Mistral says it improves accuracy while cutting the number of back and forth turns, token use, and latency compared with standard retrieval methods.
Think of ordinary document search as skimming a book's index, while Agentic Search works more like a research assistant who actually flips through the pages, cross checks footnotes, and comes back with a verified answer instead of just a page number. That multi step retrieval loop targets enterprises whose data lives across scattered file types and formats rather than one tidy database.
The launch follows Mistral's release of OCR 4 and updates to its Mistral Common library for multi image content handling and offline tokenizer support, showing the Paris based lab leaning into document and enterprise data tooling alongside its language models. Mistral continues to favor the Apache 2.0 license for many flagship releases, keeping a cost and data residency edge with European developers.
Meta's Muse Code Brings Persistent AI Teammates to Coding
Meta released a beta of Muse Code on August 5, 2026, a terminal based coding agent built by Meta Superintelligence Labs and powered by the newly released Muse Spark 1.2 model. The agent coordinates multiple persistent subagents that stay active for an entire coding session rather than resetting with every new request, and it ships with built in commands like plan, grill, and goal to structure planning and stress test proposed solutions.
In one internal test, Meta says Muse Spark 1.2 running inside Muse Code optimized Nvidia Hopper GPU kernels across more than 1,000 tool calls over sessions lasting up to 24 hours, working directly on the code rather than wrapping existing kernel libraries. Every model call, tool run, approval, and edit gets written to a local event log, letting a session be replayed exactly or resumed after a crash.
Muse Code lands directly opposite Claude Code and OpenAI's Codex, and Meta's own
published benchmarks show Claude Opus 5 still ahead on all three coding tests it compared against. Pricing starts at rates matching the earlier Muse Spark 1.1 model, alongside a cheaper contributor tier for developers willing to share their usage data with Meta.
Hugging Face and Anthropic Both Edge Toward Big Money Moves
Hugging Face has hired a bank to gauge buyer interest in a sale that could value the open model hosting platform at $13 billion or more, according to Business Insider reporting from August 24, 2026, nearly triple its $4.5 billion valuation from a 2023 funding round. Earlier this year the company turned down a $500 million investment from Nvidia that would have valued it at only $7 billion, citing concerns about one investor having outsized influence.
Separately, Anthropic is preparing to file its IPO prospectus as soon as the end of August 2026, and sources told CNBC on August 21, 2026 that the filing will name public backlash against AI and data centers as a material risk factor. A May Gallup survey found seven in ten Americans opposed new AI data centers being built near them, and some investors are reportedly projecting a valuation near $2 trillion for Anthropic once it lists on Nasdaq.
Both stories point to the same underlying tension in AI right now: the technology keeps shipping faster than public opinion can catch up with, and the companies building it increasingly have to put that friction in writing for investors. Hugging Face's talks are still early with no deal reached, while Anthropic's confidential filing from June is expected to become public within weeks.
Quick Recap
Nvidia pays Poolside about $6 billion for its Model Factory tech and 109 engineers, plus a $1 billion equity stake.
OpenAI retires o3 from ChatGPT today and cuts GPT-5.6 Sol pricing by more than 20 percent.
Google switches on Ask Gemini inside Google Chat starting today.
Grok 4.6 enters the top 10 on the Artificial Analysis Intelligence Index as SpaceXAI and Nvidia deepen ties.
Alibaba's Wan3.0 video model exits beta with 30 second clips and document inputs.
DeepSeek ships V4-Flash-Vision-Exp, adding image understanding to its fastest model.
Zhipu's GLM-5.3 nears an open weight release after strong coding and cybersecurity results.
Alibaba's Qwen3.8-Max and Qwen3.8-27B are both now open weight.
Anthropic cancels a planned Claude Sonnet 5 price increase set for September 1.
Moonshot sunsets older Kimi models on August 31 as K3 holds the open weight lead.
Mistral launches Agentic Search for navigating complex enterprise documents.
Meta's Muse Code brings persistent, auditable AI subagents to terminal coding.
Hugging Face explores a $13 billion sale while Anthropic preps an IPO naming AI backlash as a risk.
Frequently Asked Questions
What is the biggest AI news today, August 26, 2026?
The biggest story is Nvidia's roughly $6 billion deal with Poolside, which licenses Poolside's training technology and moves 109 of its engineers onto Nvidia's Nemotron model effort. It signals Nvidia is stepping further into building AI models itself rather than only selling the chips other labs use to train them.
What new AI models came out this week?
Recent releases include xAI's Grok 4.6, Alibaba's Wan3.0 video model and Qwen3.8 family, DeepSeek's V4-Flash-Vision-Exp, Zhipu's GLM-5.3, and Meta's Muse Spark 1.2 alongside its Muse Code coding agent. Most of the biggest jumps this month came from Chinese labs racing to ship open weight models.
Is Grok 4.6 open source?
No. Grok 4.6 is a proprietary model available through the xAI API, Cursor, and xAI's own Grok Build tool, and organizations need to negotiate commercial terms directly with xAI
to use it. That differs from open weight releases like Kimi K3, Qwen3.8, and the upcoming GLM-5.3 weights, which can be downloaded and self hosted.
What happened to OpenAI's o3 model?
OpenAI retired o3 from ChatGPT on August 26, 2026, after a 90 day sunset period, following the earlier retirement of GPT-4.5. The change only affects ChatGPT, so any existing conversations that used o3 automatically continue on a current model, and the API is unaffected.
Is Anthropic close to going public?
Anthropic confidentially filed paperwork to go public in June 2026 and is expected to make its S-1 prospectus public as soon as the end of August 2026, targeting a fall listing on Nasdaq. Sources say the filing will explicitly list public backlash against AI and data center construction as a risk factor, an unusually direct disclosure for a company some investors value near $2 trillion.
Recommended Blogs
Learn AI in 5 Minutes a Day
Unrot turns days like this into a five minute morning habit, breaking down model releases, funding news, and new AI tools into plain English explainers you can read over coffee. If today's roundup of Grok, GLM, and Nvidia news made sense to you, Unrot is built to keep it that way every day.
References
Nvidia pays Poolside $6 billion
DeepSeek timeline release dates
Qwen3.8-27B officially launched
Claude developer platform updates
Kimi K3 benchmarks and pricing




