New AI Models Released in August 2026: The Full List
Every new AI model released in August 2026 with dates, verified per-token prices and primary sources, plus the price changes that landed alongside them.
By Capital & Compute
Six distinct model releases landed in the first two weeks of August 2026: Qwen3.8-Max from Alibaba (August 3), Meta’s Muse Spark 1.2 (August 5) and Muse Glimmer (August 10), xAI’s Grok 4.6 (August 12), and on August 13 and 14 respectively, Google’s Gemini 3.7 Flash and Z.ai’s GLM-5.3. Every date and rate below is verified against the lab’s own page, and the price is stated as the provider states it, including when that price has an expiry date attached.
The month’s real story is not in the release list. It is that three separate labs changed what a published price means. Google’s newest Flash tier carries a rate that doubles on 1 January 2027. Anthropic cancelled a scheduled increase and made a discount permanent. DeepSeek is replacing a flat rate with peak and off-peak billing on 16 August, and the new off-peak rate is more than double the price it charges today. Prices moved in both directions in the same fortnight. This page is the August installment of a monthly series; the complete dated list lives in the record of AI model releases by month, and current per-token rates and upcoming models are in the AI model tracker.
Every AI model released in August 2026
| Model | Maker | Released | Price (in/out per Mtok) | What it is |
|---|---|---|---|---|
| Qwen3.8-Max | Alibaba | Aug 3 | $2 / $6 | 2.4T MoE flagship, open weights from Aug 13, custom licence |
| Muse Spark 1.2 | Meta | Aug 5 | $1.25 / $4.25 | Paid agentic model, 1M context, plus Muse Code terminal agent |
| Muse Glimmer | Meta | Aug 10 | Not published | 29.6B dense multimodal, Apache 2.0, runs on one consumer GPU |
| Grok 4.6 | xAI | Aug 12 | $2 / $6 | Frontier model, 500K context, cached input raised to $0.50 |
| Gemini 3.7 Flash | Aug 13 | $0.75 / $3.75 | Efficient tier, introductory rate that doubles on Jan 1, 2027 | |
| GLM-5.3 | Z.ai | Aug 14 | Not published | Coding and agent flagship on the GLM-5.2 base, weights delayed |
Five labs, and unlike July nobody quite collided. July’s pattern was three labs shipping into a single 48-hour window; August spread its releases across twelve days, then clustered at the end, with Grok 4.6, Gemini 3.7 Flash and GLM-5.3 landing on three consecutive days.
| Day of August 2026 | Releases | Models |
|---|---|---|
| 3 | 1 | Qwen3.8-Max |
| 5 | 1 | Muse Spark 1.2 |
| 10 | 1 | Muse Glimmer |
| 12 | 1 | Grok 4.6 |
| 13 | 1 | Gemini 3.7 Flash |
| 14 | 1 | GLM-5.3 |
The month a published price stopped being a single number
Three labs changed the meaning of their own rate card inside two weeks, in three different ways, and only one of those changes was a cut.
DeepSeek is the sharpest. Its API documentation states that from 16:00 UTC on August 16, 2026 it adopts peak and off-peak pricing, with off-peak set at half the peak rate, peak hours running 01:00 to 04:00 and 06:00 to 10:00 UTC. Read as a discount scheme that sounds generous. Read against the price it charges today, it is a substantial increase. DeepSeek V4 Pro currently costs $0.435 input and $0.87 output per million tokens. After the change the off-peak rate is $0.66 and $1.98, and the peak rate is $1.32 and $3.96. The cheapest hour of the new schedule is 2.3x the current output price, and the most expensive is 4.6x.
The cheapest hour of DeepSeek's new schedule costs more than twice what the same model costs today.
Google went the other way, twice, with a clock attached. Gemini 3.7 Flash launched on August 13 at $0.75 input and $3.75 output per million tokens, and Google’s Gemini API pricing page states plainly that this is introductory: on January 1, 2027 it becomes $1.50 and $7.50, with the context-cache rate going from $0.075 to $0.15. The same page now lists Gemini 3.6 Flash, released in late July at $1.50/$7.50, at the identical $0.75/$3.75. That is a 50% cut roughly three weeks after launch, and it also expires at the end of the year.
Anthropic did the rarest thing: it cancelled a planned increase. Claude Sonnet 5 launched on June 30 at $2/$10 per million tokens, described at the time as introductory pricing through August 31, 2026, reverting to $3/$15 on September 1. Anthropic’s pricing page now carries a note stating that the $2/$10 rate “is now the standard price” and that “the previously scheduled increase to $3/$15 per million input/output tokens on September 1, 2026 will not occur.” Anyone who planned a September migration off Sonnet 5 to avoid a 50% rise can stop.
| Item | Higher rate | Lower rate |
|---|---|---|
| DeepSeek V4 Flash (by hour) | $1.32 | $0.66 |
| DeepSeek V4 Pro (by hour) | $3.96 | $1.98 |
| Gemini 3.6 Flash (by date) | $7.50 | $3.75 |
| Gemini 3.7 Flash (by date) | $7.50 | $3.75 |
| Claude Sonnet 5 (cancelled) | $15.00 | $10.00 |
Qwen3.8-Max: the top open-weights score, and a licence that is not open source (August 3)
Alibaba announced Qwen3.8-Max on August 3 on the Model Studio API and published the weights on August 13 as Qwen3.8-2.4T-A95B on Hugging Face. It is a 2.4 trillion-parameter mixture of experts with 95 billion active, 512 experts with 11 activated per token, and a 262,144-token native context extensible to roughly 1.01 million. It is the first Max-class Qwen model to ship downloadable weights, and on the independent Artificial Analysis Intelligence Index it scores 58, the highest of any open-weights model on that board as of August 14.
Qwen’s own published benchmarks put it at 92.6 on GPQA Diamond, 86.6 on Terminal Bench 2.1, 67.7 on SWE-bench Pro, 56.6 on Deep-SWE 1.1 and 93.0 on PaperBench. Those are vendor-reported and have not been independently reproduced.
The licence is where the coverage has been sloppy. The model card states a custom qwen3.8-max licence, not Apache 2.0. Reading the licence text itself, two thresholds matter. Above 100 million monthly active users or US$20 million monthly revenue, the model name must be prominently displayed. More consequentially, if aggregate revenue exceeds US$50 million over any twelve consecutive months for a model-as-a-service or AI work-assistant business, a separate licence from Qwen is required before commercial use. Internal use is explicitly exempt, so a company running it on its own workloads is unaffected. A company reselling inference is not.
Muse Glimmer: Meta’s genuinely permissive release (August 10)
Meta published Muse Glimmer on Hugging Face under a plain apache-2.0 licence, with no acceptable-use addendum and no revenue threshold. It is a dense causal transformer of roughly 29.6 billion parameters including a perception encoder, with a 131,072-token context, text and image input, text output, more than 100 languages and a January 4, 2026 knowledge cutoff. It is distilled from Muse Spark and aimed at autonomous agentic work on consumer hardware: quantized to 4 bits it drops under 20 GB and fits a single 24 or 32 GB GPU.
Meta’s published benchmarks are strong: 94.7% on AIME 2026, 83.5% on GPQA Diamond, 76.0% on SWE-Bench Verified, 75.5% on MCP Atlas. Artificial Analysis scores it 35 on the Intelligence Index at high effort, which is the widest gap between vendor-reported and independent scoring of any August release, and worth weighing before adopting it on the strength of the launch table alone. There is no first-party per-token API rate, so no price appears in our registry.
Note also what did not happen: there was no Llama release in this window at all. Meta’s open-weights activity has moved to the Muse family, and coverage that still frames Meta’s roadmap around a forthcoming Llama is describing a product line the company is no longer shipping under that name.
| Primitive | Weights published | Permissive licence | Commercial use unrestricted |
|---|---|---|---|
| Muse Glimmer | Hugging Face | Apache 2.0 | No threshold |
| GLM-5.2 | Hugging Face | MIT | No threshold |
| LongCat-Flash-Lite-Sparse | Hugging Face | MIT | No threshold |
| Qwen3.8-Max | Hugging Face | Custom licence | Gated above $50M |
| Kimi K3 | Hugging Face | Custom licence | Gated above $20M |
Grok 4.6: same sticker, more expensive cache (August 12)
xAI released Grok 4.6 on August 12, 35 days after Grok 4.5. The headline rate is unchanged at $2 input and $6 output per million tokens with a 500K context window, and there is a fast variant at double the price. xAI’s stated benchmarks are an Artificial Analysis Intelligence Index of 61, GDPVal-AA v2 of 1753, CursorBench v3.2 at 69.9%, DeepSWE v1.1 at 65.9% and FrontierCode v1.1 Extended at 61.3%. Artificial Analysis independently confirms the 61 at high effort, which ties GPT-5.6 Sol and trails only Claude Opus 5 at 63 and Claude Fable 5 at 62.
The detail the launch post does not lead with sits in xAI’s model documentation: cached input rose from $0.30 per million tokens on Grok 4.5 to $0.50 on Grok 4.6. For an agent loop that replays a large cached prefix on every iteration, cached input is often the dominant line on the bill, so a 67% rise there can outweigh an unchanged sticker rate. Both models also double every rate once a prompt reaches 200K tokens, and the higher rate then applies to every token in the request, not just the ones past the threshold.
Gemini 3.7 Flash: five points of index for half the price (August 13)
Google announced Gemini 3.7 Flash on August 13 across Google AI Studio, Android Studio, the Gemini Enterprise Agent Platform and Gemini Spark for AI Pro and Ultra subscribers in more than 160 countries. Google’s own comparison against Gemini 3.6 Flash: DeepSWE v1.1 up to 65.3% from 49.0%, FrontierCode 1.1 Main to 43.6% from 34.4%, WebDev Arena Elo to 1588 from 1538, GDP.pdf to 34.0% from 22.0% and AutomationBench to 30.4% from 17.0%.
Independently, Artificial Analysis scores it 56 on the Intelligence Index at high effort. That is the number worth holding onto, because at $0.75/$3.75 it puts a model four points below Grok 4.6 and five below GPT-5.6 Sol at a fraction of their token price, at least until the end of December. Bloomberg reported the launch as arriving while Google’s top-end Gemini 3.5 Pro remains delayed, which has now been the case since Gemini 3.5 Flash shipped in May.
GLM-5.3: the same base model, post-trained harder (August 14)
Z.ai released GLM-5.3 on August 14, and it is the most interesting engineering story of the month even though it is the thinnest on published economics. It is not a new base model. Z.ai states it runs on the same base as GLM-5.2 and that every capability gain comes from scaled post-training: more task environments, more environment types and longer training. Z.ai’s model documentation confirms a 1 million-token context window and 128K max output.
The vendor-reported jumps are concentrated exactly where a post-training story would predict, on long-horizon agentic work rather than single-turn quality. Terminal-Bench 3.0 goes from 4.6 to 28.3, DeepSWE v1.1 from 46.2 to 66.9, and Agents’ Last Exam on the CLI from 23.8 to 28.5. On the security side Z.ai reports CyberGym at 84.5% from 77.2% and ExploitBench at 54.4% from 24.4%, and says the cyber capability scaled faster than it expected. All of those are Z.ai’s own numbers on Z.ai’s own runs, and none has been independently reproduced.
Two caveats matter more than the benchmark table. First, there is no published price: Z.ai’s pricing page still lists GLM-5.2 at $1.40/$4.40 per million tokens as its newest priced model and shows no GLM-5.3 rate, so this release cannot yet be compared on cost per task. Second, and unusually for Z.ai, the weights were not published at launch. GLM-5.2 shipped MIT-licensed and downloadable; GLM-5.3 is API-only for now, with Z.ai saying it will publish weights roughly two weeks after launch once safety evaluation and hardening finish, and naming no licence in advance.
That last point is the one to hold. A model that is announced as open, benchmarked as open, and covered as open, but is not yet downloadable, is the same pattern as Qwen3.8-27B and Muse Spark 1.2 this month. If the weights land on schedule with an MIT licence, GLM-5.3 becomes the strongest open coding model on these numbers. Until then it is a closed model with an open-weights promise attached.
Did Anthropic release a new model in August 2026?
No. Anthropic’s newest model remains Claude Opus 5, released July 24 at $5/$25 per million tokens, which still leads the Artificial Analysis Intelligence Index at 63 on max effort. But two things changed on the commercial side in August that matter more than a release would have.
The first is the Sonnet 5 increase being cancelled, covered above. The second is that Claude Opus 4.1 was retired on August 5, having been deprecated on June 5, though Anthropic’s pricing page still lists it as available on Amazon Bedrock and Google Cloud. Its $15/$75 rate is a useful marker of how far flagship pricing has moved: Claude Opus 5, which supersedes it and outscores it comfortably, costs a third as much. One migration trap worth flagging: temperature, top_p and top_k now return a 400 error on Claude Opus 4.7 and later, so code written against 4.1 needs editing rather than simply repointing at a new model id. Our Opus 5 pricing and benchmarks breakdown has the full picture, and the Opus 5 versus Sonnet 5 comparison covers which tier to actually run.
Did OpenAI release a new model in August 2026?
No new model. OpenAI’s most recent release remains the GPT-5.6 family (Sol, Terra and Luna), generally available since July 9, whose prices it cut on July 30, taking Luna down 80% to $0.20/$1.20 and Terra down 20% to $2/$12. Sol still sits at $5/$30 and scores 61 on the Artificial Analysis Intelligence Index at max effort, tying Grok 4.6.
The change in August was distribution rather than capability: GPT-5.6 Sol became the default model for Plus and Pro users, and Luna the default for Free and Go, as reported in early August. That matters commercially because it moves the majority of ChatGPT traffic onto models OpenAI has already made much cheaper to serve.
Which August model should you actually use?
Based on the verified rates and the independent index scores above, not the launch tables:
- Cheapest capable tier, right now: Gemini 3.7 Flash at $0.75/$3.75 with an index of 56. Put a January 1 reminder in your budget, because that rate doubles. If your workload runs past the new year, model it at $1.50/$7.50 and decide on that basis.
- Top of the board: Claude Opus 5 at 63, then Claude Fable 5 at 62, then Grok 4.6 and GPT-5.6 Sol tied at 61. Grok 4.6 is by far the cheapest of those four at $2/$6, so it is the one to price first if your work is agentic and cache-heavy, with the caveat that its cached-input rate rose.
- Self-hosting with commercial intent: Muse Glimmer if the licence matters most, since Apache 2.0 with no revenue gate is the cleanest position of any model here, and it runs on one consumer GPU. Qwen3.8-Max if capability matters most, but read the $50M clause before you build a product on it.
- Long-horizon agent work: watch GLM-5.3, but do not commit to it yet. The Terminal-Bench 3.0 jump from 4.6 to 28.3 is the largest single-benchmark move of the month, and there is no published price and no independent score to check it against. Revisit when the weights and a rate card exist.
- Anyone currently on DeepSeek: re-run your numbers before August 16. This is the one change in the month that can raise a bill without anybody touching the code.
- Anyone who planned a September migration off Claude Sonnet 5: cancel it. The increase you were avoiding is not happening.
For a head-to-head on your own workload, the model comparison tool prices two or three models against the same task, the cost-per-task calculator takes your own token profile, and the value leaderboard ranks the field by benchmark points per dollar.
The release timeline from here
What is confirmed, announced or credibly rumored as of August 14, 2026:
- Gemini 3.5 Pro (Google DeepMind): announced, still no date. It has been listed as coming soon since May, and Bloomberg framed the August 13 Flash launch as arriving while the top-end model remains delayed. Google is now the only top-three lab that has not shipped a frontier-tier model since May.
- Qwen3.8-27B (Alibaba): committed and not delivered. Alibaba pointed to the week of August 10 for the smaller sibling of Qwen3.8-Max. As of August 14 the Hugging Face page reads as an upcoming release, with no repo, model card, licence or benchmarks.
- Muse Spark 1.2 open weights (Meta): announced, unpublished. The API model shipped August 5; the weights have not appeared.
- GLM-5.3 weights (Z.ai): dated, not yet delivered. Z.ai committed to roughly two weeks after the August 14 launch, which points at the week of August 24, with no licence named. Z.ai has a good record here, since GLM-5.2 shipped MIT, but a deployment plan needs files and a licence.
- GPT-6 (OpenAI): rumor only. Still no model page, no date and no confirmation of the name.
The pattern worth reading is that five separate commitments to publish open weights are outstanding across Alibaba, Meta, Z.ai, Ant Group and Black Forest Labs. Announced is not shipped, and a release tracker that counts announcements rather than artifacts will overstate the month. The one lab that did deliver on the day, Meta with Muse Glimmer, is also the one that attached the cleanest licence.
Frequently asked questions
- What new AI models came out in August 2026?
- Six distinct models in the first two weeks: Qwen3.8-Max from Alibaba on August 3, Muse Spark 1.2 from Meta on August 5, Muse Glimmer from Meta on August 10, Grok 4.6 from xAI on August 12, Gemini 3.7 Flash from Google on August 13, and GLM-5.3 from Z.ai on August 14. Two of the six, Qwen3.8-Max and Muse Glimmer, shipped with downloadable weights on the day.
- What is the latest AI model right now?
- As of August 14, 2026 the newest release is GLM-5.3 from Z.ai, a coding and agent model that runs on the same base as GLM-5.2 with every gain from scaled post-training. It landed one day after Gemini 3.7 Flash from Google and two days after Grok 4.6 from xAI.
- Is GLM-5.3 open source?
- Not yet. Unlike GLM-5.2, which shipped MIT-licensed and downloadable, GLM-5.3 was API-only at its August 14 launch. Z.ai says it will publish the weights roughly two weeks after launch, once safety evaluation and hardening finish, and has not named a licence in advance.
- Did Anthropic release a new model in August 2026?
- No. Its newest model is still Claude Opus 5, released July 24. Anthropic did make two commercial changes in August: it cancelled the scheduled September 1 increase on Claude Sonnet 5, making the $2/$10 per million token rate permanent, and it retired Claude Opus 4.1 on the Claude API on August 5.
- Did OpenAI release a new model in August 2026?
- No. The most recent OpenAI models remain GPT-5.6 Sol, Terra and Luna, generally available since July 9 and repriced on July 30. In August OpenAI made Sol the default for Plus and Pro users and Luna the default for Free and Go.
- Is DeepSeek raising its prices?
- Yes. From 16:00 UTC on August 16, 2026 DeepSeek moves to peak and off-peak billing. For DeepSeek V4 Pro the output rate becomes $1.98 per million tokens off-peak and $3.96 at peak, against $0.87 today, so even the cheapest hour costs 2.3 times the current price. Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC.
- Is Qwen3.8-Max open source?
- The weights are downloadable, but the licence is not an open-source licence. Qwen3.8-Max ships under a custom qwen3.8-max licence rather than Apache 2.0. It requires prominent attribution above 100 million monthly active users or US$20 million monthly revenue, and requires a separate licence from Qwen once a model-as-a-service or AI work-assistant business exceeds US$50 million of revenue over twelve consecutive months. Internal use is exempt.
- What is the cheapest new AI model in August 2026?
- Gemini 3.7 Flash at $0.75 input and $3.75 output per million tokens, which scores 56 on the independent Artificial Analysis Intelligence Index. That rate is introductory: Google states it doubles to $1.50 and $7.50 on January 1, 2027.
- What AI models are coming next after August 2026?
- Gemini 3.5 Pro from Google DeepMind remains announced with no date and is now several months delayed. Alibaba committed to Qwen3.8-27B for the week of August 10 and had not shipped it as of August 14. Z.ai committed to GLM-5.3 weights roughly two weeks after its August 14 launch. Meta announced open weights for Muse Spark 1.2 that have not been published. GPT-6 remains a rumor with no model page or date.
Sources
- Alibaba Qwen (2026). Qwen3.8-2.4T-A95B model card and licence. Hugging Face. Parameters, context window, benchmarks and the custom licence terms. Verified August 14, 2026.
- Anthropic (2026). Pricing. Claude platform documentation. Claude Sonnet 5 standard rate and the cancelled September 1 increase; Opus 5 and Opus 4.1 rates. Verified August 14, 2026.
- Anthropic (2026). Model deprecations. Claude platform documentation. Claude Opus 4.1 retirement date and parameter deprecations. Verified August 14, 2026.
- DeepSeek (2026). Models and pricing. DeepSeek API documentation. Current and post-August 16 peak/off-peak rates for V4 Pro and V4 Flash. Verified August 14, 2026.
- DeepSeek (2026). News and updates. DeepSeek API documentation. V4 Pro general availability and the pricing change effective date. Verified August 14, 2026.
- Google (2026). Introducing Gemini 3.7 Flash. Google blog. Launch date, availability surfaces and vendor benchmark comparison. Verified August 14, 2026.
- Google (2026). Gemini API pricing. Google AI for Developers. Gemini 3.7 and 3.6 Flash rates and the January 1, 2027 increase. Verified August 14, 2026.
- Meta (2026). Muse-Glimmer-30B model card. Hugging Face. Apache 2.0 licence, parameters, context, knowledge cutoff and vendor benchmarks. Verified August 14, 2026.
- Meta (2026). Muse Spark. Meta developer documentation. Muse Spark 1.2 per-token rates, contributor tier and context window. Verified August 14, 2026.
- xAI (2026). Grok 4.6. xAI news. Release date, headline pricing and vendor benchmark scores. Verified August 14, 2026.
- xAI (2026). Models. xAI documentation. Cached-input rates and long-context pricing for Grok 4.5 and 4.6. Verified August 14, 2026.
- Artificial Analysis (2026). Model leaderboards. Independent model evaluations. Intelligence Index scores with reasoning-effort variants. Read August 14, 2026.
- Meituan (2026). LongCat-Flash-Lite-Sparse model card. Hugging Face. MIT licence and parameter counts. Verified August 14, 2026.
- Moonshot AI (2026). Kimi K3 licence. Hugging Face. The US$20 million revenue threshold that triggers a separate commercial agreement. Verified August 14, 2026.
- Z.ai (2026). GLM-5.2 model card. Hugging Face. MIT licence, 753B parameters and 1M context. Verified August 14, 2026.
- Z.ai (2026). GLM-5.3. Z.ai model documentation. Context window and max output. Verified August 14, 2026.
- Z.ai (2026). Pricing. Z.ai documentation. Confirms GLM-5.2 at $1.40/$4.40 is the newest priced model and that no GLM-5.3 rate is published. Verified August 14, 2026.
- Asif Razzaq (2026, August 14). Z.ai ships GLM-5.3 without retraining the base model. MarkTechPost (secondary). Release date and the vendor-reported benchmark deltas against GLM-5.2, which Z.ai’s own client-rendered blog post could not be read directly to confirm. Verified August 14, 2026.
- Capital & Compute. New AI models released in July 2026. The previous installment of this series.