Top AI News Today: August 23, 2026 (13 Biggest Stories)

August 23, 2026 brought a mix of price cuts, a surprising cybersecurity story, and a handful of new open weight coding models. The biggest story is Z.ai's GLM-5.3, an open coding model that found over a thousand real security bugs in widely used software, but this was also a week where OpenAI and Google both cut prices on their flagship models and Meta and Alibaba pushed out new open weight releases.

Below are the 13 biggest AI stories from today, covering new models, what those models can actually do, and the business and policy moves shaping the industry around them. Model releases and updates make up most of this list, since that is usually what readers searching for AI news want to know first, followed by a look at new capabilities and a couple of industry moves worth tracking.

Z.ai releases GLM-5.3, a coding model with surprise cyber skills

Z.ai released GLM-5.3 on August 14, 2026, an update to its GLM-5.2 coding model that reuses the exact same underlying base model, with every improvement coming from extra training after the fact, a step called post-training. The company reports a 50 percent jump on its own internal Code Bench and says GLM-5.3 leads open source models on Terminal-Bench 3.0, scoring 28.3 percent versus 4.6 percent for GLM-5.2. The model is live now through the GLM Coding Plan, starting at $18 a month, and through Z.ai's ZCode tool, but the full API and downloadable weights are being held back for what the company calls safety hardening.

The bigger story is what GLM-5.3 can do in cybersecurity. Working with outside security teams, Z.ai says the model found 2,436 vulnerabilities across 269 real software projects, including 1,097 rated medium to high severity, in systems like the Linux kernel, the WebKit browser engine, and FreeBSD. Think of it like hiring an extremely fast code reviewer who can read a million lines of software before lunch and flag the weak spots a human might take months to notice. Z.ai says the model's hacking skill grew faster than expected during training, which is why open weights are delayed by roughly two weeks instead of shipping right away.

GLM-5.3 is the first release in the GLM series to gate open weights behind extra review. Independent testers note the model still trails Anthropic's Fable 5 and OpenAI's GPT-5.6 Sol on several of the hardest coding benchmarks, and reports say it already flagged a security issue in Cursor, the coding tool recently bought by SpaceX. Watch for the public weights, which Z.ai has pointed to around the end of August, since that is when self hosted, budget conscious teams will finally get a chance to run this generation of the model on their own hardware.


OpenAI cuts GPT-5.6 Sol's price by more than 20 percent

OpenAI cut the price of its flagship GPT-5.6 Sol model by more than 20 percent on August 21, 2026, part of a promotion running at least through November 21. Sol now costs $4 per million input tokens and $20 per million output tokens, down from $5 and $30. Sol is the top tier in OpenAI's three-model GPT-5.6 family, alongside the mid-range Terra and budget Luna, and it powers coding, research, and agent work in ChatGPT and the API.

The discount matters because Sol has been positioned as OpenAI's answer to Anthropic's frontier models, and price is now as much a battleground as raw capability. A cheaper flagship means a business running large coding or research workloads through Sol pays less for the same quality of answers, a bit like an airline dropping business class fares to fill more seats without changing the service.

The cut follows OpenAI's July move to slash Luna's price by 80 percent and Terra's by 20 percent, showing a pattern of frequent price adjustments rather than one time launches. It also lands the same week Google cut Gemini 3.7 Flash pricing and DeepSeek raised its own API prices, a reminder that the cost of frontier AI is moving in different directions across providers right now, with some labs chasing volume through discounts and others charging more once a model proves itself in wide use.


Google launches Gemini 3.7 Flash at half the price

Google released Gemini 3.7 Flash on August 13, 2026, calling it its most intelligent workhorse model yet for coding and AI agents. The model costs $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026, half of what Gemini 3.6 Flash cost at launch just three weeks earlier. On Google's own benchmarks, it scored 43.6 percent on FrontierCode 1.1 Main, up from 34.4 percent, and 65.3 percent on the DeepSWE v1.1 coding test, up from 49.0 percent.

Gemini 3.7 Flash targets everyday development work rather than the hardest reasoning problems, the kind of workhorse job a mid-size delivery van handles compared with a specialized freight truck. Google says it generates more complete, working code in fewer attempts and sticks closer to design references like screenshots. It now also

powers Gemini Spark, Google's personal AI agent, for subscribers in more than 160 countries.

The cheaper price roughly doubles on January 1, 2027, rising to $1.50 and $7.50 per million tokens, so the current rate is a limited window rather than a permanent cut. The release also comes as Google's larger Gemini 3.5 Pro flagship remains delayed with no new timeline, leaving February's Gemini 3.1 Pro as the company's newest large reasoning model.


DeepSeek's V4-Pro goes fully live, and prices jump

DeepSeek moved its V4-Pro model out of preview and into general release on August 13, 2026, four months after first showing it in April. The production build, called V4-Pro-0813, scored 87.9 on Terminal Bench 2.1 and 62.7 on the DeepSWE benchmark, sharp jumps from the earlier preview version. It runs on a reported 1.6 trillion parameter design with a 1 million token context window.

The launch arrived with a steep price increase. DeepSeek introduced peak and off-peak billing, and output tokens during busy hours now cost $3.96 per million, up from a flat $0.87 before, an increase the company describes as up to 1,100 percent on some token types. Even after the hike, DeepSeek remains far cheaper than most Western rivals, similar to a discount airline that still undercuts full-service carriers even after raising a few fares.

The move suggests DeepSeek, long known for near-free pricing, is starting to charge closer to what serving a frontier-class model actually costs, especially since it has no cloud business of its own to subsidize token prices. DeepSeek is separately reported to be raising close to $8 billion in new funding at roughly a $74 billion valuation, a sign that investors still see room to grow even as the company adjusts its pricing strategy.


Alibaba open-sources its 2.4 trillion parameter Qwen3.8-Max

Alibaba released its largest model yet, Qwen3.8-Max, in early August, and this month followed through on a promise to open its weights, publishing them on Hugging Face

around August 12 through 14. The model has 2.4 trillion total parameters but activates only 95 billion of them per request, and supports a 1 million token context window across text, images, and video.

The open weight release matters because Alibaba had kept its most recent flagship Qwen models closed, so this marks its first time releasing a Max class model publicly. A companion smaller model, Qwen3.8-27B, also shipped under an Apache 2.0 license and runs on a single ordinary GPU, similar to how a factory sells both an industrial machine and a compact home version of the same tool. Alibaba reports Qwen3.8-Max ranks fifth in Text Arena and second in Vision Arena, trailing mainly Anthropic's Claude line.

One catch: the openly downloadable 2.4 trillion parameter checkpoint is text only, without the vision or full 1 million token context the hosted version offers, so developers wanting the complete multimodal package still need Alibaba's paid service at $2 per million input tokens and $6 per million output tokens. Alibaba's shares rose on both the initial announcement and the open weight release, a sign investors are watching the open model race as closely as developers are.


Meta ships Muse Spark 1.2 and its first coding agent

Meta released Muse Spark 1.2 and its first terminal based coding agent, Muse Code, on August 5, 2026. Spark 1.2 is a coding focused update to July's Muse Spark 1.1, trained in part using the earlier model to generate and grade its own practice coding tasks, a kind of self improvement loop. Pricing stays at $1.25 per million input tokens and $4.25 per million output tokens.

Muse Code can run multiple background helper agents at once inside a coding session, so one part of the system keeps working on a task while another checks results, somewhat like a construction crew where different workers handle framing and inspection at the same time. On Meta's own tests, Spark 1.2 still trails Claude Opus 5, scoring 82.9 percent against Opus 5's 86.7 percent on Terminal-Bench 2.1.

This is Meta's third Muse Spark release in four months, a fast pace for a company that only entered the paid AI model business in July. Meta is leaning on aggressive pricing rather than raw benchmark wins as its main pitch against Anthropic and OpenAI in the coding tools market, betting that developers will choose a cheaper, slightly less capable model over a pricier frontier option for everyday work.


Meta open-sources Muse Glimmer for offline coding

Meta released Muse Glimmer on August 10, 2026, a smaller 30 billion parameter open weight model built for local, always on coding agents. It is a dense model rather than a mixture of experts design, distilled from the larger Muse Spark, and ships under the permissive Apache 2.0 license.

The headline feature is that Muse Glimmer runs entirely offline on a single ordinary 24 gigabyte consumer graphics card, which matters for developers who want an AI coding helper without sending code to a cloud server, similar to keeping a reference book on your desk instead of calling a library every time you need a fact. It pairs a language model with a dedicated image and screenshot understanding component.

Glimmer arrives alongside a wave of open, GPU friendly agent models this month, including Alibaba's Qwen3.8-27B, as labs compete to put capable coding assistants directly on developer laptops rather than only through paid cloud APIs.


Kimi K3 is splitting into two membership plans

Moonshot AI's Kimi K3 chatbot is preparing to split its subscription plans, with a banner on kimi.com as of August 20, 2026 warning that new membership tiers are coming that separate general use from coding focused access, though the company says current subscribers will not be affected.

Kimi K3, a 2.8 trillion parameter open weight model launched in July, remains the largest open weight AI system released to date and briefly paused new signups last month after demand overwhelmed Moonshot's computing capacity. Splitting plans into a general Kimi Membership and a separate Kimi Code Membership lets the company match its limited graphics card capacity more precisely to how people actually use the model, the same logic an airport uses when it opens separate lines for carry on only and checked bag passengers.

The change reflects a broader pattern among fast growing Chinese AI labs this year: ship a headline grabbing open model, then scramble to manage the compute needed to serve everyone who wants it. Moonshot is separately reported to be preparing for a Hong Kong stock listing.


OpenAI previews an ultrafast version of GPT-5.6 Sol

OpenAI began previewing an Ultrafast version of GPT-5.6 Sol this month, a speed optimized mode the company says runs up to 14 times faster than the standard model by using specialized Cerebras hardware. Access remains limited to a select group of customers while OpenAI studies how the extra speed changes real products.

The idea is to make Sol usable for tasks that need answers in a split second, such as voice conversations, live customer support, or financial research, where waiting several seconds for a reply breaks the experience, much like the difference between a phone call with normal delay and one with an awkward satellite lag. OpenAI says its own staff have used Ultrafast to analyze system logs during live incidents and to run several research passes in a single workday that previously took overnight.

Businesses can join a waitlist by describing their workload and latency needs. Speed focused variants like this are becoming their own category alongside raw intelligence gains, following a similar pattern to Google's low latency Flash tier and Anthropic's effort dial approach on Opus 5.


Z.ai launches OpenVuln, a scanner built on GLM-5.3

Alongside GLM-5.3, Z.ai launched OpenVuln, a scanning tool that uses the new model to search code repositories for security weaknesses. Vulnerability scanning is normally slow, careful work, and pairing it with a strong coding model compresses how long it takes to comb through a large codebase for weak points.

Z.ai is rolling OpenVuln out to selected trusted security partners first, with wider access planned within two weeks as part of the same staged release used for GLM-5.3 itself. The company frames the tool as helping defenders, comparing it to giving every security team a much faster flashlight to search a dark building for problems before an intruder finds them first.

The launch lands in a year when AI labs are increasingly public about the double edged nature of strong coding models: the same skill that finds and fixes a bug can, in the

wrong hands, help someone exploit it. Z.ai's public disclosure ledger tracks every vulnerability its model finds through to a fix rather than letting findings sit unaddressed.


Anthropic turns on Auto Mode by default in Claude Code

Anthropic began switching on Auto Mode by default for Claude Code, its coding agent, for Pro, Max, and Team accounts starting August 14, 2026. Auto Mode lets Claude carry out multi step coding tasks with less back and forth approval from the person using it.

The change reflects a broader shift toward AI coding agents that work more independently. Instead of asking permission at every small step, the tool acts more like an experienced junior developer who checks in occasionally rather than one who asks before typing every line. Anthropic reports its safety classifier catches a large share of risky commands before they run.

The update follows months of steady feature additions to Claude Code, including new sandboxing rules for file access and cross session messaging, as coding agents from Anthropic, OpenAI, and Meta all move toward longer, less supervised task runs this year.


Cognition AI is in talks for a $40 billion valuation

Cognition AI, maker of the Devin coding assistant, is in early talks to raise new funding that could value the company at more than $40 billion, according to a Bloomberg report from August 12, 2026, up over 50 percent from the $26 billion valuation it held less than three months earlier.

The company's revenue is reportedly approaching a $1 billion annual run rate, roughly double what it was at its last funding round, a growth pace that helps explain why investors are circling again so soon. It is a reminder that the AI coding tools market, not just the underlying model labs, is drawing enormous investor interest right now.

Cognition is one of several coding focused AI companies attracting outsized valuations this year, alongside broader AI funding that industry trackers estimate topped $407 billion globally in just the first half of 2026, more than all of 2025 combined.


Anthropic starts watermarking Claude's output worldwide

Anthropic began adding invisible, machine readable watermarks to text and files produced by Claude models released after August 2, 2026, to comply with the European Union's AI Act. Older models are set to get the same treatment by December 2, 2026.

The watermark is built into the text itself, so it can survive copying and pasting elsewhere, but Anthropic says it can be stripped by resaving, reformatting, or taking a screenshot, so it is not a foolproof way to detect AI writing. The rule applies worldwide rather than only to European users, similar to how a food label requirement in one country sometimes ends up printed on every box a company ships anywhere.

The EU's transparency requirement took effect August 2, 2026 and carries fines up to 15 million euros or 3 percent of a company's global revenue for non compliance, pushing AI providers toward similar labeling systems regardless of where their users are located.


Quick Recap

Z.ai released GLM-5.3, a coding model that also found over 1,000 real security bugs, and delayed open weights for safety review.

OpenAI cut GPT-5.6 Sol's price by more than 20 percent through at least November 21.

Google launched Gemini 3.7 Flash at half the price of its predecessor, good through the end of 2026.

DeepSeek's V4-Pro left preview and API prices rose sharply, especially at peak hours.

Alibaba open sourced its 2.4 trillion parameter Qwen3.8-Max, its first open Max class model.

Meta shipped Muse Spark 1.2 and its first coding agent, Muse Code.

Meta open sourced Muse Glimmer, a 30 billion parameter model that runs on one consumer GPU.

Kimi K3 is splitting into separate general and coding membership plans.

OpenAI is previewing an Ultrafast GPT-5.6 Sol that runs up to 14 times faster for select customers.

Z.ai launched OpenVuln, a vulnerability scanner built on GLM-5.3.

Anthropic turned on Auto Mode by default in Claude Code for Pro, Max, and Team accounts.

Cognition AI is in talks for a funding round that could value it above $40 billion.

Anthropic started watermarking Claude's output worldwide to meet EU AI Act rules.


Frequently Asked Questions

What is the top AI news today, August 23, 2026?

The biggest story is Z.ai's release of GLM-5.3, an open coding model that also found over 1,000 real security bugs in software like the Linux kernel, alongside price cuts from OpenAI and Google and a fresh wave of open weight releases from Meta and Alibaba.

What new AI model was released today?

No single frontier model launched on August 23 itself, but several major releases from earlier in August, including GLM-5.3, Gemini 3.7 Flash, DeepSeek V4-Pro, and Meta's Muse Spark 1.2, are still shaping today's pricing, access, and safety news as companies roll them out further.

Why did GLM-5.3 delay its open weights?

Z.ai says the model's cybersecurity skill grew faster than expected during training, well beyond what the company originally planned for, so it is holding back public weights for about two weeks of extra safety review and hardening before release.

Is GPT-5.6 Sol cheaper now?

Yes. OpenAI cut Sol's price by more than 20 percent on August 21, 2026, to $4 per million input tokens and $20 per million output tokens, and the company says this promotional pricing will hold through at least November 21, 2026.

What is Qwen3.8-Max and is it open source?

Qwen3.8-Max is Alibaba's largest model, with 2.4 trillion total parameters and 95 billion

activated per request. Alibaba published open weights for it in mid August, though the openly downloadable version is text only, unlike the full hosted version, which also handles images and video.

Why did DeepSeek raise its API prices?

DeepSeek moved V4-Pro from a limited preview to a full production release and introduced peak and off-peak billing at the same time, which pushed some token prices up sharply even though the model remains cheaper than most Western competitors.

Is Kimi K3 still free to use?

Yes, Kimi K3 remains free to try on kimi.com with usage limits, but Moonshot AI is splitting its paid membership plans into separate general and coding tiers, a change that does not affect existing subscribers.

Recommended Blogs

How to Use Claude AI

How to Use Google Gemini

ChatGPT Free for Beginners 2026

Best AI Tools for Coding 2026

What Is Agentic AI

Learn AI in 5 Minutes a Day

Unrot turns days like this one into a five minute lesson, breaking down new models, pricing shifts, and what they mean for your work without the jargon. If today's roundup felt like a lot to track, that is exactly the kind of AI news Unrot summarizes every day.

References

Z.ai launches GLM-5.3

GLM-5.3 found bug in Cursor

GPT-5.6 Sol pricing update

Google introduces Gemini 3.7 Flash

DeepSeek V4-Pro official launch

Alibaba launches Qwen3.8-Max

Meta introduces Muse Code

Muse Glimmer open weight model

Moonshot pauses Kimi K3 signups

OpenAI previews Ultrafast Sol

Anthropic turns on Claude Code Auto Mode

Cognition AI funding talks

You might also like...

Deepen your knowledge in ai news

Explore all stories →