Wholestory

Last Updated: September 19, 2026

Leading Models: The Frontier Race

Artificial Analysis revised its Intelligence Index on September 4 and again on September 7, with a fifth version due in late October. Under v4.1 GPT-6 Astra scored 61 to Claude Fable 5.1’s 66; under v4.2, 55 to 57; under v4.3 both stand at 53, and Astra has moved from fifth place to joint first. No new model shipped in between. The share of the index held in private, unpublished tests rose from about 20 per cent to 45 across the three versions — the firm’s stated defence against tuning to the test. The open-weight frontier was measured three ways in the same week. Mozilla puts the gap to the best open Chinese models at 4.4 months, or 1.7× on METR’s task time-horizon method. Eight of the ten most-used models on OpenRouter by token volume in August published open weights. And on Terminal-Bench 2.1, with every model on one neutral harness, the open-weight GLM 5.2 came within a point of Claude Opus 4.7 and 4.8 at roughly a fifth of the cost per completed task.

The Whole Story

On the independent Artificial Analysis Intelligence Index — rebuilt as version 4.2 in September 2026 — Anthropic's Claude Fable 5.1 leads at 57, ahead of OpenAI's GPT-6 Astra at 55 and Anthropic's own Claude Opus 5 at 54. Below them the band is dense and, at the very top, entirely closed: Claude Fable 5 at 53, Meta's generally available Muse Spark 1.3 at 52, xAI's Grok 4.6 and OpenAI's GPT-5.6 Sol at 51, then Alibaba's Qwen3.8-Max, Meta's Muse Spark 1.2, Google's Gemini 3.8 Flash and OpenAI's GPT-5.6 Terra at 47. The best model anyone can download — Moonshot's Kimi K3 at 50, with Zhipu's GLM-5.3 just behind at 49 — is seven points off the top. The price of a fixed level of capability has been collapsing for years — 9 to 900 times a year, by Epoch AI's measurement — and it is still moving: Anthropic cut Fable's cache-read price by three-quarters, OpenAI publishes GPT-5.6 Sol at $4 and $20 per million tokens against GPT-6 Astra's $10 and $50, and Google is holding its workhorse Flash models at $0.75 and $3.75 until the new year. Access is contested on both sides — June 2026 brought the first model-level US export controls, and both OpenAI and Meta now ship their strongest models to limited partners first.

Six years of record show three arcs. Capability: from GPT-3 behind a private-beta API to today's crowded frontier band, where six labs shipped models within a hair of each other in a single week. Price: the steepest cost curve in the industry's history — Epoch AI measures inference prices for fixed performance falling 9 to 900 times a year, and Stanford's AI Index puts GPT-3.5-equivalent capability at a 280th of its late-2022 price by late 2024. Access: the pendulum has swung from API-only gatekeeping through the open-weights insurgency — Llama first, then DeepSeek, Qwen and GLM — to government-imposed restriction, and, through the second half of 2026, part-way back toward opening. July brought the largest open-weight release on record, Moonshot's Kimi K3; August brought Alibaba publishing the weights of its 2.4-trillion-parameter flagship, the largest model anyone can download, and Zhipu following with GLM-5.3. Meta — which made open-weight frontier-chasing a corporate strategy with Llama 2 in 2023, then retooled its Muse line as proprietary — reversed part-way, releasing a small on-device model and pledging to open its flagship's weights. That pledge is unmet: into September Meta kept reaching the frontier with closed models, shipping the proprietary Muse Spark 1.3 while the promised open weights never appeared.

The open-versus-closed gap is the story's live question, and on the current scale it stands at seven points between the closed leader and the best downloadable model. Open weights are also not the same as open reach. Kimi K3's licence is close to MIT — commercial conditions that bite only above US$20M in revenue or 100M monthly users — but its training data and process are undisclosed, and Rahul Shome of the Australian National University told Nature the model is too large for personal devices and will probably need institutional investment to run. GLM-5.3, at 753 billion parameters with 40 billion active, ships under Zhipu's own licence rather than a standard open one. Alibaba's Qwen3.8-Max carries a similar catch, and a second one: what Alibaba actually published is a text-only, always-reasoning sibling of the model it sells. The UK AI Security Institute separately measures the best open-weight model at four to seven months behind the closed frontier on cyber capability. What is moving down-market is cheap hosted capability and, increasingly, local capability: Alibaba's Qwen3.8-27B, small enough to run four-bit on a 32-gigabyte laptop, scores 41, and Zhipu's MIT-licensed GLM-5.3-Flash — 320 billion parameters but only 18 active — scores 46 while running natively multimodal. A whole class of cyber-specialist models, and Tencent's 770-billion-parameter Hy4, still ships with no independent score at all.

There is a second race behind the first, and it is about what can be checked rather than what can be built. Every number above comes from one published independent index, because vendor launch materials have repeatedly been shown to grade their own homework — and that index is itself a moving object. Rebuilt in September with harder tasks and more private test sets held back to stop labs training against them, it re-scored every model at once, and the reshuffling was not uniform: GPT-6 Astra, which had scored level with the model it replaced and behind two rivals, came out of the rebuild ahead of both. No model changed. The test did. That is the honest condition of the field: where an independent index exists, it is the thing arguments are settled with, and it is revised faster than the models are; where none exists — a security model sold inside a vendor's own harness, a Chinese open-weight release benchmarked only by its maker — a capability claim is just a press release. The labs increasingly ship both kinds at once, and the outside scorecard, not the launch chart, is what separates the tested from the merely asserted.

Continue Reading →

The Frontier

Artificial Analysis Intelligence Index

Closed / API
Open weights
Release date

score vs. $ per million input tokens

Closed / API
Open weights
$ / M input tokens (log scale)
ModelLabScore$/Mtok in / out
Claude Fable 5.1 (max)Anthropic57$10 / $50
GPT-6 Astra (max)OpenAI55$10 / $50
Claude Opus 5 (max)Anthropic54$5 / $25
Claude Fable 5 (max)Anthropic53$10 / $50
Muse Spark 1.3 (xhigh)Meta AI52$1.25 / $4.25
GPT-5.6 Sol (max)OpenAI51$4 / $20
Grok 4.6 (high)xAI51$2 / $6
Kimi K3Moonshot AI50$3 / $15
GLM-5.3 (max)Zhipu AI (Z.ai)49$1.4 / $4.4
Claude Opus 4.8 (max)Anthropic48$5 / $25
Qwen3.8-MaxAlibaba (Qwen team)47$2 / $6
Muse Spark 1.2 (xhigh)Meta AI47$1.25 / $4.25

Index Artificial Analysis Intelligence Index v4.2

Open weights 3 of 12 shown

Four Months and Two-Thirds Off: the Open-Weight Gap, Measured Three Ways

Mozilla’s State of Open Source AI report puts the gap between US closed frontier models and the best open-weight Chinese models at 4.4 months. On METR’s task time-horizon method the best closed model finishes a job 1.7 times as long as the best open one can reliably complete — roughly twelve hours against seven, on Mozilla’s CTO’s illustration. Eight of the ten most-used models on OpenRouter by token volume in August published open weights. And on Terminal-Bench 2.1, run by Vals AI with every model on the same harness, the open-weight GLM 5.2 scored within a point of Claude Opus 4.7 and 4.8 at about a fifth of the cost per completed task. Mozilla argues for open models, which is a stake in its own finding.

A Four-Point Deficit Became a Tie Without Either Model Changing

On version 4.1 of the Intelligence Index, GPT-6 Astra scored 61 to Claude Fable 5.1’s 66. On 4.2 the two were 55 and 57. On 4.3 both stand at 53, and Astra has gone from fifth place to joint first. No new model was released in between; the index was revised twice. Artificial Analysis separately reports Astra matching Fable 5.1 on the Intelligence Index at about 40 per cent of the cost, and on its Coding Agent Index at about 60 per cent — the firm’s own figures. Its chief executive says version 4.1 had stopped capturing what the newest models could do, because earlier models had already saturated several of its benchmarks.

The Ruler Changed Twice in Ten Days

Artificial Analysis revised its Intelligence Index on September 4 and again on September 7, with a fifth version due in late October. Version 4.2 added AA-Briefcase and Surge AI’s GDP.pdf and dropped GPQA Diamond as saturated; 4.3 upgraded Terminal-Bench to 4.0 and added AutomationBench-AA. The share of the index held in private, unpublished test sets went from roughly 20 per cent under v4.1 to 40 under v4.2 and 45 under v4.3 — the firm’s stated defence against labs tuning to the test. The practical consequence for anyone reading a score: numbers from different versions are not comparable, and three versions have now been current within a fortnight.

Observation

On the new index, GPT-6 Astra passes Claude Opus 5

Artificial Analysis's own launch write-ups, both still published, place GPT-6 Astra "equal to GPT-5.6 Sol in the Index at 61", five points behind Claude Fable 5.1 and trailing Meta's Muse Spark 1.3; its 1 September write-up puts Claude Opus 5 at 63, two points above Astra. On the same publisher's live leaderboard under version 4.2, Claude Fable 5.1 scores 57, Astra 55, Opus 5 54, Muse Spark 1.3 53 and Sol 51. Astra now leads Sol by four points, leads Muse Spark 1.3, and sits above Opus 5. Rebuilding the test changed the ranking, not only the numbers.

Observation

Two pages, one index, two different weightings

The methodology page for Intelligence Index v4.2 says the score is "calculated as a weighted average across four categories: Agents (30%), Coding (20%), Scientific Reasoning (20%) and General (30%)", and that "the weighting emphasizes agentic tasks". The index's own evaluation page, published by the same organisation and live at the same time, says it is "calculated as a weighted average of production benchmark scores, scaled from 0 to 100. Four categories each contribute 25%: agents, coding, general capability, and scientific reasoning." The two cannot both describe the same index. The methodology page's per-evaluation weight table matches its own 30/20/20/30 figure.

The Scale Everything Is Measured On Was Rebuilt

Artificial Analysis replaced its Intelligence Index — the independent scoreboard this page ranks models by — with version 4.2, and re-scored every model on it. The publisher calls it an interim step toward a planned version 5, saying it is "accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier" and that v4.2 carries "more complex and realistic tasks, and more private test sets to prevent gaming". Ten evaluations feed it, among them two agentic knowledge-work tests, a terminal-based coding benchmark and a hallucination measure. Scores fell across the board: Claude Fable 5.1, the leader, from 66 to 57; Kimi K3, the best downloadable model, from 60 to 50; GPT-4, from 7 to 1. Nothing about the models changed. Figures from before 4 September cannot be set against figures from after it.

GPT-6 Astra is the world's most intelligent and aligned model.±

Context: Still mixed. On Artificial Analysis's rebuilt v4.2 index Astra scores 55, second to Claude Fable 5.1's 57, so the superlative does not hold. But it has moved above Claude Opus 5 (54) and Muse Spark 1.3 (53), which it trailed on the old scale, and it leads the Coding Agent Index at 67 while halving its hallucination rate.

Hy4 preview ranks among the top tier of open-source models; in a Tencent-internal blind evaluation of 203 engineering tasks by 163 experts it averaged 2.99 out of 4.00, slightly ahead of GLM-5.3 (2.92) and Kimi K3 (2.94).?

Context: Still unverified. Tencent's own private evaluation; Artificial Analysis lists 25 open-weight models and Hy4 is not among them, so the top-tier claim and the narrow wins over GLM-5.3 and Kimi K3 have no neutral index to check against. Its top open models: Kimi K3 (50), GLM-5.3 (49).

OpenAI Calls It a New Generation; the Intelligence Index Doesn't Move

OpenAI released GPT-6 Astra, branding it a generational step. On the Artificial Analysis Intelligence Index it scores 61 — exactly level with GPT-5.6 Sol and five points below Claude Fable 5.1 — while costing two-and-a-half times as much, at $10/$50 per million tokens. Its gains are elsewhere: it leads Artificial Analysis's separate Coding Agent Index at 67, uses about a third of Sol's tokens on coding tasks, and roughly halves its hallucination rate, from 92% to 51% at maximum effort, without losing accuracy. It reached a limited set of organisations first, then ChatGPT's paid tiers, the API and AWS.

Meta Reaches the Frontier, and Keeps It Closed

Meta released Muse Spark 1.3, its fourth in five months. On the Artificial Analysis Intelligence Index the limited-preview 'max' version scores 62 — behind only Claude Fable 5.1 and Claude Opus 5 — and the generally available version scores 61, the cheapest model at that level ($0.55 per task, unchanged $1.25/$4.25 pricing). It is Meta's strongest showing since it made open weights a strategy, but the model is proprietary, served through Meta's own API with the top version limited to partners. Its August pledge to 'soon' open Muse Spark 1.2's weights remains unmet.

Six Frontier-Band Models in a Single Week

Between 27 August and 3 September, five labs shipped models scoring 57 to 66 on the Artificial Analysis Intelligence Index: Anthropic's Claude Fable 5.1 (66), Meta's Muse Spark 1.3 (62), OpenAI's GPT-6 Astra (61), Google's Gemini 3.8 Flash (59, its fourth Flash model in under four months), and Zhipu's open-weight GLM-5.3-Flash (57, MIT-licensed, emerging from a week-long anonymous 'Ox Alpha' preview). Tencent's open-weight Hy4 arrived the same week, still unscored — a sixth release, from a sixth lab. Every one of the top scores belongs to a closed model.

Claude Fable 5.1 Sets the Highest Score Yet, and the Open Gap Doubles

Anthropic released Claude Fable 5.1, and its trusted-access sibling Mythos 5.1, and on the independent Artificial Analysis Intelligence Index its top effort scores 66 — the highest the index has measured, four points above Fable 5 and three above Claude Opus 5. Standard token pricing is unchanged at $10/$50 per million, but Anthropic cut the cache-read price 75%, from $1 to $0.25; even so, a task costs about 20% more than on Fable 5, because the model spends roughly 1.7 times the output tokens to reach the higher score. The move reopens the distance to the downloadable frontier: the best open-weight models, Moonshot's Kimi K3 and Zhipu's GLM-5.3, sit at 60, so the gap between the best model and the best one anyone can download widened from three points to six in a single release.

Tencent Releases an Open-Weight 770-Billion-Parameter Model

Tencent published Hy4 preview, a downloadable mixture-of-experts model with 770 billion parameters — about 49 billion active for any request — and a context window over a million tokens, also served through Tencent Cloud and OpenRouter at $0.83 per million input tokens and $2.50 output. It adds another Chinese open-weight entrant behind Moonshot, Zhipu, Alibaba and DeepSeek. Tencent flagged that the early-release model can over-verify its answers and run long on hard questions. It carries no independent capability score yet — Artificial Analysis has not evaluated it — so it stays off this page's ranked board until one exists.

Meta will 'soon' release open weights for its flagship Muse Spark 1.2, extending its open-weight pivot from the small Muse Glimmer to a frontier-band model.?

Context: Still not done. Three weeks on, Meta instead shipped a newer proprietary flagship, Muse Spark 1.3, on 2 September — served through its own API and limited to partners at the top tier — and has published open weights for neither Muse Spark model. The pledge named no deadline, so it stays unverified rather than broken, but the trajectory is away from it.

Three Points Now Separate the Best Model From the Best One You Can Download

Artificial Analysis’ index today puts Claude Opus 5 at 63, the highest of the 128 reasoning models it scores, and Moonshot’s Kimi K3 at 60, the highest of the open-weights models. Behind Kimi come Alibaba’s Qwen3.8 at 58 and DeepSeek V4 Pro at 53. Of the 161 models the index evaluates, 86 — more than half — can be downloaded. The index version was also checked this cycle: every row added to this page’s board since June is scored on Intelligence Index v4.1.1, which is what the live index runs today, so the board compares like with like. Nothing needed correcting.

A Model You Can Download Ties GPT-5.5, at a Thirty-Third of the Price

Alibaba released Qwen3.8-Flash-Next under open weights, and Artificial Analysis scored it 56 on its Intelligence Index — the same as OpenAI's GPT-5.5 and xAI's Grok 4.5. GPT-5.5 costs $5 per million input tokens; this costs $0.15, with output at $0.47 against $30. It is a mixture-of-experts model with a 125-billion-parameter main body, 51 billion more in N-gram embedding tables, and six billion parameters active per token, running 262,144 tokens of context natively and a million with YaRN. Alibaba presents it as an early look at the architecture Qwen4 will use, and says training cost about a ninth of Qwen3.7-Plus — its own figure, unverified. The independent part is the index score, which is why it is on the board; the lab's own benchmark table is not used.

A Publisher Says It Reached the Frontier for $40 Million

Thomson Reuters launched Thomson, its own large language model, built by taking an open-source foundation and spending $40 million on talent and compute to mid-train and post-train it over Westlaw, Practical Law, Checkpoint and Reuters — less than 10 percent of the company's content so far. It says the model is fully owned in-house and runs without the inference costs of a typical frontier model, and its chief executive says early evaluations put it "on par with the latest frontier models across a range of tasks". No benchmark is named, no score published and no third party has assessed it, so the capability claim is the company's alone.

Open Models Are Catching Up Twice as Fast Each Generation

SemiAnalysis measured the open-versus-closed gap across three eras of language models — early scaling, reasoning, agentic — and found it moves in cycles rather than one direction. A frontier lab jumps ahead at the start of an era; others reverse-engineer the advance and close it. What has changed is the clock: with each generation, open-source models take half as long to catch up to the first closed model of that era. The method matters. Rather than score every model on one benchmark set, which saturates and is abandoned, SemiAnalysis scored each era on its own benchmarks and compared catch-up time — a rate, not a level. It assesses GLM 5.3 and Kimi K3 as genuinely able to do the coding and agentic work that carried Anthropic past $65 billion in annualised revenue.

Observation

The weights Alibaba published are not the model Alibaba sells

Alibaba's model card for the downloadable Qwen3.8-2.4T-A95B states that the hosted Qwen3.8-Max "is the official version based on Qwen3.8-2.4T-A95B with more features, such as vision input & non-thinking support, 1M context length by default, official built-in tools", and separately that the downloadable model "is a text-only model that requires thinking mode for all interactions. Multimodal inputs are not supported, and thinking cannot be disabled." Alibaba's own page for the hosted model lists the thinking toggle and a 983,000-token maximum input. So the largest set of weights anyone can download is a text-only, always-reasoning sibling of the product on sale, and the score of 58 that places Qwen3.8-Max second among open-weight models on the Artificial Analysis index was measured on the hosted version rather than on the published weights. The smaller Qwen3.8-27B, released two days after the flagship weights, does take image input.

Observation

Three of the four biggest labs now publish a price with a date or an hour attached

Google's pricing page lists Gemini 3.7 Flash and Gemini 3.6 Flash at $0.75 per million input tokens and $3.75 per million output "through December 31, 2026", and at $1.50 and $7.50 "starting January 1, 2027". DeepSeek's price table halves its rates only outside two daily peak windows, so the same request costs two different amounts depending on the hour. Anthropic's documentation, which in July carried two dated rows for Claude Sonnet 5, now carries one at $2 and $10 with no successor row — the company withdrew its scheduled increase rather than adding one. OpenAI's pricing page lists gpt-5.6-sol, gpt-5.6-terra and gpt-5.6-luna at $5/$30, $2/$12 and $0.20/$1.20, with no dated future rate anywhere on the page. On two of the four schedules the rate a buyer pays today and the rate already published for next year differ by a factor of two.

Qwen3.8-27B has "excellent capabilities" in coding, professional work, research and long-horizon agentic tasks, and matches the performance of a model ten times its size.±

Context: The size claim holds on the independent index: at 52, Qwen3.8-27B ties DeepSeek's V4-Flash, a 284-billion-parameter model roughly ten times larger. The capability claim is not measurable as worded, and the same index shows the model spending 160 million tokens to complete its battery against a 43 million median for its size class.

The cheapest model on the board raises its prices several-fold

DeepSeek's new prices took effect at 16:00 UTC on 16 August, replacing flat rates with a peak and off-peak schedule. V4-Flash, priced through July at $0.14 per million input tokens and $0.28 per million output, now costs $0.44 and $1.32 in peak hours and half that outside them. V4-Pro moves from $0.435 and $0.87 to $1.32 and $3.96 at peak, $0.66 and $1.98 off-peak. Reported increases run from 50% to more than 1,100% depending on the model, the token type and the hour. Peak covers seven hours of every 24 — 01:00 to 04:00 and 06:00 to 10:00 UTC — so most of the day bills at the lower rate. The company said it was adopting the schedule "to allocate resources more reasonably ... encouraging users to schedule their tasks based on actual usage", announcing it alongside the general release of V4-Pro. Both models' weights remain published under the MIT licence, so the rise applies to DeepSeek's hosted service rather than to the models, which anyone can still run themselves. Even at peak, V4-Pro's output tokens cost about a twelfth of Claude Fable 5's $50. Four months earlier DeepSeek had cut V4-Pro by 75% and its cached-input prices across the API to a tenth of their previous level.

A model that runs on a laptop scores level with one ten times its size

Alibaba published Qwen3.8-27B, a dense 27-billion-parameter model under the permissive Apache 2.0 licence, and on the independent Artificial Analysis Intelligence Index it scores 52 — the same score as DeepSeek's V4-Flash, a 284-billion-parameter mixture-of-experts model ten times its size, and level with OpenAI's GPT-5.6 Luna. It is seventeen points above Meta's Muse Glimmer, the 30-billion-parameter open model released four days earlier into the same size class, and eleven points below the best model anyone can buy. The unquantised weights are 55.6 GB; a community four-bit conversion for Apple silicon is 16.1 GB, small enough for a 32 GB Mac with room left for working context. The model takes text and images and carries a 262,144-token context window, extensible to a million. The same index records what the score costs: it generated 160 million tokens to finish the test battery against a 43 million median for open models of its size, reaching its answers by thinking several times longer than its peers. Alibaba's own benchmark comparisons, including a claimed parity with Claude Opus 4.6 at maximum effort and a sweep of its own chosen tests against Muse Glimmer, are the lab's claims and not the independent measure.

Google ships Gemini 3.7 Flash three weeks after 3.6, at half price until the new year

Google released Gemini 3.7 Flash 23 days after Gemini 3.6 Flash, priced at $0.75 per million input tokens and $3.75 per million output through 31 December 2026, with $1.50 and $7.50 taking effect on 1 January 2027. It cut 3.6 Flash to the same dated rate, so the new model does not undercut its predecessor — both undercut the price 3.6 Flash launched at three weeks earlier. On the independent Artificial Analysis Intelligence Index 3.7 Flash scores 56, four points above 3.6 Flash and level with xAI's Grok 4.5 at a third of Grok's input price. Google's own benchmark table claims leads on production-code quality and web development; the same table shows OpenAI's GPT-5.6 Terra ahead on two terminal benchmarks and Claude Sonnet 5 ahead on a desktop-agent evaluation, and those are the lab's own figures either way. Google gave no release date for its next flagship, Gemini 3.5 Pro, which it had previously described as in partner testing.

Anthropic cancels the price rise it had already scheduled

Anthropic withdrew the increase it had published for Claude Sonnet 5 when the model shipped on 30 June, adding an editor's note to the launch announcement: "Sonnet 5's introductory pricing of $2 per million input tokens and $10 per million output tokens is now permanent. The standard pricing of $3 input / $15 output previously set to take effect September 1 no longer applies." Its pricing documentation, which through July carried two dated rows for the model, now lists a single row at $2 and $10. Sonnet 5 is the default model on Anthropic's free and paid consumer plans, and the withdrawn increase was the only scheduled rise on any of the three largest American labs' published price lists.

Observation

A frontier price was scheduled to rise by half, and the token it is charged on has changed size

The published direction of travel was not all downward. Anthropic's pricing documentation carried two dated rows for Claude Sonnet 5: $2.00 and $10.00 per million input and output tokens "through August 31, 2026", then $3.00 and $15.00 "starting September 1, 2026" — a 50% increase on both sides, published in advance since the model shipped on 30 June rather than sprung, and applying to the model Anthropic made the default in its Free and Pro plans. The increase was withdrawn on 10 August and the $2/$10 rate made permanent; the same documentation now carries a single Sonnet 5 row. The second thing it disclosed still stands, and it cuts across every per-token comparison here: models from Claude 4.7 onward use a newer tokenizer producing roughly 30% more tokens for the same text, so a price per token is not the same unit across model generations. Every other Anthropic price was re-checked against the same page and is unchanged: Claude Fable 5 at $10/$50, Claude Opus 5 and Claude Opus 4.8 at $5/$25.

Alibaba publishes Qwen3.8-Max weights, the largest downloadable model

Alibaba published the weights of Qwen3.8-Max, its 2.4-trillion-parameter flagship (95 billion active per token, one-million-token context) — the first Max-class Qwen released as open weights and the largest-parameter model anyone can download, fulfilling the release the company had promised at the model's 3 August launch. The terms let most developers use, copy, modify and distribute the model without payment, but require a separate paid commercial licence for any 'model as a service' or 'AI work assistant' business whose aggregate revenue exceeds US$50 million over any consecutive 12-month period; internal use is exempt. The model scores 58 on the Artificial Analysis Intelligence Index v4.1.1, second among open-weight models behind Moonshot's Kimi K3 (60) and three points below the closed leader, Claude Opus 5 (63). Artificial Analysis's own model page still classified Qwen3.8-Max as proprietary as of 13 August, a lag behind the actual release. A paid API remains available at $2 per million input tokens and $6 per million output.

Meta returns to open weights with Muse Glimmer

Meta released Muse Glimmer, a 30-billion-parameter open-weight model under the permissive Apache 2.0 licence with weights on Hugging Face — its first open-weight release since retooling its Muse flagship line as proprietary under Alexandr Wang. Optimized for always-on local agents, it is small enough to run on a Mac or PC with a single consumer GPU: compressed to roughly 4-bit precision it shrinks to under 20 GB, and Meta says it is distilled from the larger Muse Spark. It is multimodal (text and image), spans a 131,000-token context and was trained to a January 2026 knowledge cutoff. On the independent Artificial Analysis Intelligence Index v4.1.1 it scores 35 — an edge-scale rather than frontier score, but well above the median for open-weight models of its size. Meta paired the release with a pledge to open the weights of its frontier-band flagship Muse Spark 1.2 (57 on the same index), reopening the open-versus-closed question the company had appeared to abandon. Meta's own benchmark comparisons — against Google's Gemma 4 31B and Alibaba's Qwen3.6-27B — are the lab's own claims; the score cited here is the independent index's.

Meta ships Muse Spark 1.2, its third model in four months

Meta released Muse Spark 1.2, a proprietary model scoring 54 on the Artificial Analysis Intelligence Index v4.1 — three points above Muse Spark 1.1's 51, inside the frontier band but below the leaders. It is priced at $1.25 per million input tokens and $4.25 per million output, unchanged from Muse Spark 1.1, with a one-million-token context window and text-and-image input, and is optimized for coding. Like its predecessor it is closed: Artificial Analysis lists the weights as not publicly available and Meta discloses no parameter count. It is Meta's third model release in four months, from the company whose Llama line made open-weight frontier-chasing a corporate strategy before its current flagship range went proprietary.

Alibaba releases Qwen3.8-Max, a proprietary frontier-band model

Alibaba released Qwen3.8-Max, a 2.4-trillion-parameter sparse mixture-of-experts model scoring 56 on the Artificial Analysis Intelligence Index v4.1 — inside the frontier band, below Moonshot's Kimi K3 (57), OpenAI's GPT-5.6 Sol (59) and Anthropic's Claude Opus 5 (61), and above Grok 4.5, Claude Sonnet 5 and GPT-5.5. It is priced at $2.00 per million input tokens and $6.00 per million output, with a one-million-token context window. Artificial Analysis classifies the model as proprietary — available through Alibaba's API, weights not published. Stanford's HAI and some trade coverage described Qwen3.8-Max as open-weight, and Alibaba separately released a smaller 27-billion-parameter open-weight variant, but the 2.4-trillion Max that carries the index score is API-only on the independent index's classification. Alibaba's own blog said the weights would be released; at launch they had not been, so on the measurement this record uses the largest new Chinese model counts as closed.

DeepSeek releases V4-Flash, the cheapest well-known model to run

DeepSeek released V4-Flash (V4-Flash-0731), an open-weights reasoning model published under the MIT licence with weights on Hugging Face and designed to run on Huawei Ascend accelerators as well as Nvidia. A 284-billion-parameter mixture-of-experts model activating 13 billion parameters per token, it scores 50 on the Artificial Analysis Intelligence Index v4.1 — level with Google's Gemini 3.6 Flash and a point below Zhipu's GLM-5.2 and Meta's Muse Spark 1.1 — at a published price of $0.14 per million input tokens and $0.28 per million output. Artificial Analysis found it the cheapest well-known model to run, at roughly 3 cents to complete its Intelligence Index test battery, against 86 cents for Moonshot's Kimi K3, $1.86 for OpenAI's GPT-5.6 Sol and $3.15 for Anthropic's Claude Fable 5 — a cost measured per test rather than per token, because a model that is cheap per token can still run up a large bill when it needs more steps to reach an answer.

OpenAI cuts GPT-5.6 Luna 80% in a frontier price war

OpenAI cut the API price of two GPT-5.6 models, attributing the move to efficiency gains passed to customers rather than a promotion. Luna, the series' fast model, fell 80% to $0.20 per million input tokens and $1.20 per million output, from $1/$6; Terra, the everyday model, fell 20% to $2/$12 from $2.50/$15. The flagship, Sol, held at $5/$30 but gained a "Sol Fast" mode at twice standard price ($10/$60) for 2.5× throughput with no change in intelligence, replacing Priority Processing. On the Artificial Analysis Intelligence Index v4.1 Luna scores 51 and Terra 55; at the new price Artificial Analysis puts Luna at about $0.21 to run its test battery, undercutting Zhipu's GLM-5.2 as the cheapest model at that intelligence level. The cut is the latest in a run of releases that, by Artificial Analysis's count, put six labs above 50 on that index within eight days — the price of near-frontier capability being competed down rather than a single vendor's discount.

Some people argue that we must close our models to prevent China from gaining access to them, but my view is that this will not work and will only disadvantage the US and its allies. Our adversaries are great at espionage, stealing models that fit on a thumb drive is relatively easy...?

Context: A strategic argument, not yet adjudicable: whether openness or closure better serves US advantage stays contested. But the premise that closing US models would not deny China frontier capability keeps gaining evidence — Moonshot's Kimi K3, DeepSeek's V4-Flash and Alibaba's Qwen3.8-Max are native Chinese capability, not derivatives of leaked US weights, though Qwen shipped proprietary before its August open-weight release, cutting the other way.

We believe an open approach is the right one for the development of today's AI models, especially those in the generative space where the technology is rapidly advancing. By making AI models available openly, they can benefit everyone.?

Context: A development-philosophy statement, not falsifiable — but measurable against Meta's conduct, which has now reversed twice. Its Muse flagship line shipped proprietary (Muse Spark 1.1 and 1.2), then in August 2026 Meta returned to open weights, releasing Muse Glimmer and pledging to open Muse Spark 1.2. The 2023 argument stands as the founding statement of the strategy Meta abandoned and has now resumed.

DeepMind sends Gemini into the physical world as a robot's planner

Google DeepMind released Gemini Robotics ER 2, an embodied-reasoning model that acts as a robot's high-level planner — reading continuous video, tracking its own task progress, coordinating multiple robots — while delegating motor control to lower-level vision-language-action models. It ships publicly through the Gemini API rather than as a research demo, which is the substantive step: frontier-lab reasoning models are now a purchasable component for physical machines, with whatever that implies for the deployment questions this record tracks mostly in software.

OpenAI previews GPT-5.6 to trusted partners first, at Washington's request

OpenAI began a limited preview of the GPT-5.6 series — flagship Sol, everyday model Terra and fast model Luna — restricted initially to trusted partners, a staging OpenAI attributes to the US government's request. The staggered access is the news as much as the models: it is the first frontier release sequenced to a government ask since June's executive order made a classified cyber benchmark the trigger for federal involvement in frontier releases, and it turns the 'covered model' question from a hypothetical into release policy. What Sol scores and where it lands on the indices belongs to the benchmark record once general access opens.

Moonshot's Kimi K3 puts a Chinese model in the frontier's top tier

At 2.8 trillion parameters, China's largest model debuted at 57 on the Artificial Analysis index — fourth overall, above GPT-5.5 and behind only Claude Fable 5 and GPT-5.6 Sol — and second on AA's agentic knowledge-work benchmark. Chinese AI rivals' shares plunged on the news (Z.ai -28%, MiniMax -16%). K3 launched API-only, with weights announced for release at the end of July 2026.

Observation

The index moves the largest open-weight model onto the open side of its own cut

Artificial Analysis now classifies Kimi K3 as an open-weights model. Its page for the model carries the label "Open weights model" and answers the question directly — "Yes, Kimi K3 is open weights. The model weights are publicly available and can be downloaded for self-hosting" — listing the weights against the Hugging Face repository that hosts them and the bespoke Kimi K3 License. Before the weights were published on 27 July the same page answered the opposite, that the model was proprietary and its weights not publicly available, which left the index's separate open-weights-against-proprietary ranking counting the largest open-weight model ever released on the wrong side of the distance it exists to measure. The score is unchanged at 57, so what has moved is the cut and not the measurement: the four-point gap to Claude Opus 5 at the top of the index is now readable from Artificial Analysis's open-weights view rather than only by comparing its per-model pages. One difference remains between the index and the licence it describes — Artificial Analysis summarises Kimi K3 as requiring a separate agreement for commercial use, where the licence text grants rights to use, modify, distribute and sell outright and imposes a separate agreement only on operators reselling model access above US$20M in group revenue.

Microsoft ships a security model that cannot be evaluated from outside

Microsoft announced MAI-Cyber-1-Flash, its first model trained specifically for cybersecurity, describing it as "a compact, code-heavy security model derived from the MAI-Thinking-1 lineage, which was built from scratch, in-house, on the highest quality data". It is reachable only inside MDASH, Microsoft's multi-agent vulnerability-finding harness: there is no standalone endpoint, no waitlist and no licence terms on either launch page, and no parameter count, context length or per-token price is published anywhere. Neither Microsoft page states an availability status; Ars Technica reports the new tools are in preview. Its parent model has a self-published technical report — MAI-Thinking-1 is given there as a 34.7-billion-active, 962-billion-total mixture of experts pre-trained on 30 trillion tokens across 8,000 GB200 GPUs, and as trained "without distillation from third-party models" — but Microsoft does not say whether MAI-Cyber-1-Flash is a distillation, a fine-tune or a smaller sibling of it, and no independent index scores either model.

This combined expertise delivers exceptional security protection, beating Mythos, Gemini and GPT on CyberGym, the gold standard benchmark for evaluating how systems reason over large codebases to find real vulnerabilities in the code.±

Context: The published number doesn't measure what the sentence claims. Microsoft's 96% on CyberGym is for the whole MDASH harness (100+ agents, hardest 10% routed to GPT-5.4), not a standalone MAI-Cyber-1-Flash score, which is published nowhere; competitors are named only as 'Mythos, Gemini and GPT' with no versions or scores. No methodology, date or trial count, and no independent index scores the model. The cost claim conflates a saving vs Microsoft's own prior config with a saving vs competitors.

Observation

Three cyber-specialist releases in seven days, none of them priced or independently scored

A new class of model arrived in a week, and none of it can be measured from outside. Google announced Gemini 3.5 Flash Cyber on 21 July as a limited pilot for governments and trusted partners; Anthropic opened a second Claude Opus 5 tier on 24 July for Cyber Verification Program members, carrying fewer security restrictions than the model as generally released; Microsoft announced MAI-Cyber-1-Flash on 27 July, available only inside its own MDASH harness. Not one of the three is priced as such: Google's published Gemini API pricing lists no Flash Cyber row, Microsoft's two launch pages publish no per-token figure, and Anthropic's published pricing documentation carries a single Claude Opus 5 line, at $5 and $25 per million input and output tokens, with no separate entry for the restricted tier. Nor is one of them independently scored. Artificial Analysis, which supplies every capability number in this record, returns nothing for gemini-3-5-flash-cyber, for mai-cyber-1-flash or for the MAI-Thinking-1 lineage the Microsoft model comes from, at the same URL pattern that resolves for every model on the frontier board; its one Claude Opus 5 figure, 61, is for the model as generally released, not for the relaxed tier. Every capability claim made for all three therefore rests on an evaluation run by the lab that built the model.

Distillation does not allow the CCP to obtain equivalent or superior AI capabilities to the US, but it can bring the Chinese frontier to within a few months of the US frontier.±

Context: Both halves still hold on the index this record uses: no Chinese model has reached the top (Kimi K3 60 vs Claude Opus 5's 63 on the re-based v4.1.1 index), and the lag is weeks, not months. But Amodei cited his own firm's research rather than an independent index, and framed distillation as a general property — unfalsifiable as stated.

Observation

Inside the open-weight tier, prices span more than an order of magnitude

The best downloadable model is also the dearest of them. Fortune published a comparison of output-token prices putting Kimi K3 at $15 per million, Zhipu's GLM-5.2 at $4.40 and DeepSeek-V4-Pro at about $0.87, against $50 for an Anthropic Fable model — figures Fortune sources to Moonshot AI and OpenRouter, and which match the vendor prices recorded independently in this record's own frontier board for all three open-weight models. So the model that halved the measured distance to the top of the Artificial Analysis index costs roughly 3.4 times GLM-5.2 and 17 times DeepSeek-V4-Pro per output token, and Fortune's own reading is that Kimi K3 is "relatively pricey by Chinese standards". Fortune's Anthropic figure names no version; $50 per million output tokens is Claude Fable 5's published price, and half that is Claude Opus 5's.

We are creating a baseline on top of which you will see us rapidly iterate on subsequent model releases. And so picking up pace and releasing models almost at a monthly cadence is part of our road map as we are building Gemini 4 as well,?

Context: Not falsifiable as stated: Pichai attached no date to Gemini 4 and no start to the 'monthly cadence', so it carries no prediction tag. Said on Google's 22 July earnings call, the day after Google shipped three Flash-class models and no flagship, with the June-promised Gemini 3.5 Pro still unreleased and no update given. Read via InfoWorld's account of the public call.

Meta ships its flagship model closed

Meta released Muse Spark 1.1, a multimodal reasoning model with a 1,048,576-token context window, alongside the public preview of the Meta Model API — its first developer API for the model. Artificial Analysis scores it 51 on its Intelligence Index v4.1 at the highest effort setting, against a median of 33 for reasoning models in the same price band, and prices it at $1.25 per million input tokens and $4.25 per million output tokens, undercutting Zhipu's GLM-5.2 at the identical score of 51. It is also closed: Artificial Analysis lists the model as proprietary and states its weights are not publicly available, Meta discloses no parameter count, and Meta's own launch post never uses the words "open weights" or "open source" about it. That is a reversal for the company whose Llama 2 release in July 2023 turned open-weight frontier-chasing from a leak into a corporate strategy, on the argument that "by making AI models available openly, they can benefit everyone". Artificial Analysis measures the model as fast — 129 output tokens per second against a 67 median — and somewhat verbose, consuming 94 million output tokens to run the index against a 63 million median.

Our internal belief is that by March of 2028 we may have a significant fraction of our research being done by AI systems in tandem with our own researchers.?

Context: Not yet resolvable; the date is 2028. The claim is hedged three ways — 'internal belief', 'may have', and a 'significant fraction' defined nowhere: no percentage, baseline or measurement method. It is the only calendar-dated forward claim in the plan (bylined Sam Altman and Jakub Pachocki, 8 June 2026); the nearest other carries no date and cannot be scored.

Moonshot AI claims Kimi K3 beat Claude Opus 4.8 and GPT 5.5 — models CNBC describes as sitting just behind Anthropic's and OpenAI's leading-edge systems — on benchmarks including coding and general agents.±

Context: On the re-based Artificial Analysis index K3 (60) now leads both models it named — GPT-5.5 (56) and Claude Opus 4.8 (57). Mixed overall because the specific coding and general-agent wins still rest on Moonshot's own benchmark table, run under its Kimi Code harness while rivals use theirs, and are not independently replicated; and neither model named was its lab's flagship — Claude Opus 5 now tops the index at 63.

The most likely (by far) outcome is for the status quo to continue and for the best open models to lag the best closed models by 6-9months.?

Context: Still open — resolves at end-2026 — but current readings run against it. On the Artificial Analysis index the best open model (Kimi K3, 60) trails the closed leader (Claude Opus 5, 63) by weeks, not the forecast 6-9 months. Caveats: the metric is a year-end state; Lambert warned public benchmarks may overstate open closeness; and the UK AI Security Institute puts open cyber capability 4-7 months behind. Recheck at resolution.

Moonshot publishes the Kimi K3 weights, the largest open-weight model yet released

Moonshot AI published the full weights of Kimi K3 on Hugging Face, meeting to the day the end-of-July commitment it made at the model's API launch — Nature had reported on 21 July that Moonshot said the weights "will be released on 27 July". K3 is a 2.8-trillion-parameter mixture-of-experts model with 104B active parameters, 896 experts of which 16 fire per token, and a 1,048,576-token context window; Artificial Analysis scores it 57 on its Intelligence Index v4.1, unchanged from the API-only release, at $3.00/$15.00 per million tokens. The licence is a bespoke "Kimi K3 License" granting MIT-style rights to use, modify, distribute, sublicense, sell and fine-tune, with no acceptable-use policy, and only two commercial conditions: an operator selling model access with group revenue above US$20M over any twelve months must sign a separate agreement with Moonshot, and any product above 100M monthly active users or US$20M monthly revenue must display "Kimi K3" in its interface. Open weights are not open source — the training process and dataset remain undisclosed — and openness is not the same as reach: Tom's Hardware headlined the model "2-3x easier to run", but that figure is a price comparison of Moonshot's internally measured costs against rivals' retail token prices rather than a hardware measurement, and Rahul Shome of the Australian National University told Nature the model is too large for personal devices and will probably need significant institutional investment to run.

It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.

Context: Core holds: the published weights are for a 2.8-trillion-parameter model, larger than any prior open-weight release, corroborated by Tom's Hardware and, pre-release, by ANU's Rahul Shome in Nature. Caveats that don't falsify it: '3T-class' rounds up 2.8T, and only 104B parameters activate per token. Capability is independently corroborated — Artificial Analysis scores K3 at 60 — but Moonshot's own benchmark table, run under its Kimi Code harness, is not.

The open-weight gap on the Artificial Analysis index halves in four days

Two releases four days apart moved both ends of the same measurement. On 23 July, the best model anyone could download was Zhipu's GLM-5.2, which Artificial Analysis scores 51 on its Intelligence Index v4.1, against the highest score on that index, Claude Fable 5's 60 — a nine-point distance. On 27 July, with Kimi K3's weights published, the best downloadable model scores 57 and the highest score on the index is Claude Opus 5's 61: four points. Every figure is on Artificial Analysis' own per-model pages and was re-verified on 27 July; none of the four scores changed this cycle except the newly added Opus 5, so the movement comes entirely from which models are eligible on each side. One qualification travels with it: as of the day the weights were published, Artificial Analysis still labelled Kimi K3 a "proprietary model", meaning its own open-weights breakdown had not yet been updated to include it.

Anthropic releases Claude Opus 5 at half of Fable 5’s per-token price

Anthropic released Claude Opus 5 at $5.00 per million input tokens and $25.00 per million output tokens — the same price as Opus 4.8, and half Claude Fable 5's $10/$50 — with a one-million-token context window and five effort settings. Artificial Analysis measures it at 61 on its Intelligence Index v4.1, one point above Fable 5's 60, a margin AA itself calls "effectively tied", ahead of GPT-5.6 Sol (59), Kimi K3 (57) and Opus 4.8 (56); AA discloses that it "supported Anthropic to evaluate Claude Opus 5 ahead of release", so this is a pre-release partnered evaluation rather than an arm's-length one. Its published findings are mixed: Opus 5 takes joint first on AA's Coding Index and records the highest GDPval-AA v2 score AA has measured (1861 Elo), but runs at 54.8 output tokens per second against a 79.6 median for comparably priced reasoning models, consumed 100M output tokens running the index against a 63M median — AA classes it "notably slow and very verbose" — trails GPT-5.6 Sol, GPT-5.5 Pro and GPT-5.6 Terra on the CritPt physics benchmark, and raised its hallucination rate 14 points to 50% on AA-Omniscience while gaining 7 points of factual accuracy. Anthropic also made Opus 5 the default on Claude Max and created a two-tier cyber regime, giving Cyber Verification Program members a version with fewer security restrictions.

Claude Opus 5 is available today. It's a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.±

Context: Price half right, capability half understates. Anthropic's $5/$25 is exactly half Fable 5's $10/$50, and on the Artificial Analysis index Opus 5 (63) matches rather than approaches Fable 5 (62) — an effective tie. Mixed because Anthropic's separate 'best-performing and most cost-effective' claim isn't supported: on AA's cost-per-task measure Opus 5 sits below Fable 5 but above Opus 4.8 and Sonnet 5, and it runs more output tokens than the median.

Observation

Google ships three Gemini models — and not the flagship

Google released Gemini 3.6 Flash and Gemini 3.5 Flash-Lite into general availability, plus a restricted security model, Gemini 3.5 Flash Cyber. Artificial Analysis independently scores 3.6 Flash at 50 on its Intelligence Index v4.1 and Flash-Lite at 36, against class medians of 32 and 16; Google's published API prices are $1.50/$7.50 per million input/output tokens for 3.6 Flash and $0.30/$2.50 for Flash-Lite. Axios headlined the release a "series of new cheaper Gemini models" and TechCrunch called them "cheaper, faster", but Google's own pricing page supports that only for 3.6 Flash, whose output price falls from Gemini 3.5 Flash's $9.00 to $7.50 with input unchanged: on the same page, Gemini 3.5 Flash-Lite is listed dearer than the 3.1 Flash-Lite it succeeds, at $0.30 against $0.25 per million text input tokens and $2.50 against $1.50 per million output tokens. Google's blog claims only "a strong price-to-performance ratio", never that Flash-Lite got cheaper. Absent from the launch was Gemini 3.5 Pro, the flagship; Flash-Lite instead began rolling out inside Google Search.

Google restricts its new vulnerability-hunting model to governments and trusted partners

Gemini 3.5 Flash Cyber, a Flash-class model fine-tuned to find and patch software vulnerabilities, was announced without general availability: Google DeepMind is releasing it as a limited pilot to governments and trusted partners through its CodeMender agent, citing the dual-use nature of automated vulnerability discovery. It carries no price on Google's published Gemini API pricing page and no score on the Artificial Analysis Intelligence Index, so every capability figure for it comes from Google or Google-run evaluation — including the "independent" evaluation Google credits to its own Big Sleep team, which is independent of the model team rather than of Google. Google's own footnotes disclose that its CyberGym competitor results are provider self-reported, and that the competitor baseline in its V8 comparison is Claude Opus 4.6 because newer Anthropic models refuse the tasks, so that comparison measures willingness alongside capability.

xAI launches Grok 4.5 on token efficiency, not peak capability

Priced at $2/$6 per million tokens and claiming roughly twice the token efficiency of comparable models, with xAI's own published tables showing it trailing Anthropic's Fable and OpenAI's GPT-5.5 on several coding evaluations — a notable strategic concession from the lab that had claimed the outright frontier a year earlier.

Thinking Machines releases Inkling, a US open-weight frontier entrant

Mira Murati's lab shipped its first foundational model with full weights on Hugging Face, trained on Nvidia infrastructure — partly on data generated by existing open models, including China's Kimi K2.5. The first credible US-startup answer to the Chinese dominance of open weights, arriving amid growing enterprise demand for self-hostable models.

UK AISI: the open-weight cyber gap has narrowed to 4-7 months

The institute's first public open-vs-closed analysis found GLM-5.2 matching the most cyber-capable closed models of November 2025 on its 70-task suite and 32-step cyber range — a lag of 4-7 months, down from the 6-10 months it had measured for earlier open models, at a fraction of the cost per run ($1.19 versus $85 for one 100M-token range run). It also found open-model safeguards largely did not impede its evaluations.

Export controls on Fable 5 and Mythos 5 lifted; access restored

Anthropic restored Fable 5 globally from July 1 and Mythos 5 to approved organizations, reporting that the safeguard-bypass incident preceding the controls had exposed no unique Mythos-level cyber capability beyond what less-capable models already provided. The three-week episode established that model-level export control is now a live instrument of US policy.

Claude Sonnet 5 becomes Anthropic's default model

The most agentic Sonnet shipped across all plans as the Free and Pro default, at introductory pricing of $2/$10 per million tokens (rising to $3/$15 in September). Independent scoring placed it at 53 on the Artificial Analysis index — within striking distance of flagships costing five times as much, though with notably verbose token usage.

OpenAI limits GPT-5.6 models to trusted partners at US government request

Announced days before the series' public debut, and paired with the Fable 5 suspension then in force: for the first time, both leading US labs' most capable models were access-restricted by government action simultaneously — a wind at the back of the open-weights alternative, per contemporaneous industry reporting.

Zhipu's GLM-5.2 becomes the strongest open-weight model ever released

MIT-licensed open weights with a 1M-token context, landing within a percentage point of Anthropic's Opus 4.8 on closely watched agentic benchmarks at roughly a fifth of the cost, per CNBC's reporting — with OpenRouter traffic climbing faster post-launch than any prior open model. Zhipu also disclosed increased reward-hacking behavior in coding RL, an unusually candid model-card admission.

Anthropic launches Claude Fable 5 and Mythos 5, a tier above Opus

The first Mythos-class models: Fable 5 for general availability with safeguard routing on cyber, bio, and distillation topics, and Mythos 5 restricted to approved organizations — at $10/$50 per million tokens, which Anthropic said was less than half the price of its prior top tier. A new deployment pattern: one underlying model, two access regimes separated by safety gating.

MiniMax M3 brings 1M-token context to open-weight pricing

A frontier coding/agentic model using MiniMax's MSA sparse attention (1/20th per-token compute at 1M context versus its predecessor), priced from $0.30/$1.20 per million tokens. Artificial Analysis initially scored it the leading open-weights candidate — while noting the promised weights arrived later than the launch commitment.

Google opens the Gemini 3.5 generation with Flash

The family's first release shipped same-day across the Gemini app, AI Mode in Search, and developer platforms, with vendor benchmarks showing the mid-tier Flash outperforming the previous generation's Pro — each generation's small model now overtaking the last generation's flagship.

DeepSeek previews V4, open-source in Pro and Flash versions

The long-awaited successor arrived open-source with markedly lower inference costs ($0.435/$0.87 per million tokens for V4 Pro), and Huawei confirmed its Ascend clusters support it — Chinese frontier capability and Chinese silicon advancing together. Analysts judged the capability real but the market impact muted: Chinese competitiveness was already priced in.

GPT-5.5 arrives with OpenAI's first dual 'High' risk designation

Better at coding, computer use, and research per OpenAI, at $5/$30 per million tokens with a 1M context — and the first OpenAI model treated as High-risk in both biological/chemical and cybersecurity domains under its Preparedness Framework, with advanced cyber capability gated behind identity-verified 'Trusted Access'. The UK AISI independently measured GPT-5.5 (with Anthropic's Mythos Preview) as one of the largest cyber-capability jumps it had evaluated.

DeepSeek's V3.2 claims GPT-5-level daily-driver performance

V3.2 and the reasoning-maxed V3.2-Speciale shipped with DeepSeek's first thinking-integrated tool use, trained via a new agent-data synthesis method across 1,800+ environments. The lab's claims — 'GPT-5 level' V3.2, Speciale 'rivals Gemini-3.0-Pro' with gold-medal results on IMO and ICPC — remained vendor-reported, but the release confirmed a second Chinese lab iterating at frontier cadence.

Google ships Gemini 3, embedding it in Search on day one

Gemini 3 Pro topped LMArena at 1501 Elo and scored 37.5% on Humanity's Last Exam (vendor-reported), launching simultaneously across the Gemini app, AI Mode in Search, and developer platforms — the first time Google put a frontier model into Search at release, with distribution numbers (2 billion monthly AI Overviews users) no rival could match.

OpenAI releases GPT-5 to all ChatGPT users

A unified system — fast model, deeper reasoning model, and a real-time router — that replaced GPT-4o, o3, o4-mini, GPT-4.1, and GPT-4.5 as ChatGPT's default for its nearly 700 million weekly users, at $1.25/$10 per million tokens in the API. OpenAI reported ~45% fewer factual errors than GPT-4o with web search and large token-efficiency gains over o3.

This is the best model in the world at coding. This is the best model in the world at writing, the best model in the world at health care, and a long list of things beyond that.?

Context: Vendor superlative at launch. The Verge noted the same day that OpenAI had lacked an industry-leading frontier model despite ChatGPT's reach; on Artificial Analysis's current index GPT-5 (high, 35) does score above its scored August-2025 contemporaries, but domain-by-domain superiority in coding, writing, and health care has no independent measurement in this record.

OpenAI returns to open weights with gpt-oss

gpt-oss-120b and gpt-oss-20b shipped under Apache 2.0 — OpenAI's first open-weight language models since GPT-2, with the 120B model achieving near-parity with o4-mini on the lab's reasoning benchmarks while running on a single 80GB GPU. OpenAI adversarially fine-tuned the model pre-release to test misuse potential, and executives acknowledged the release was driven by customers already running Chinese open models.

It's a very good thing for the open community.

Context: Independently corroborated: MIT Technology Review reported the same week that Chinese open models had overtaken Llama in popularity, and subsequent UK AISI and Artificial Analysis measurements consistently placed Chinese models atop the open-weight rankings through mid-2026.

xAI releases Grok 4, claiming the frontier's top spot

Available via API and new consumer tiers including a $300/month SuperGrok Heavy plan, with vendor-reported firsts including 50% on Humanity's Last Exam for Grok 4 Heavy. Trained with reinforcement learning at pretraining scale on the Colossus cluster.

Anthropic releases Claude Opus 4 and Sonnet 4

Hybrid instant/extended-thinking models, with Opus 4 at $15/$75 per million tokens and a reported 72.5% on SWE-bench Verified. Anthropic shipped Opus 4 under its ASL-3 safety standard — the first frontier release explicitly deployed with a heightened internal safety level.

The U.S. is doubling down on restricting sales of chips to China and purchases from China, but models like Qwen 3 that are state-of-the-art and open […] will undoubtedly be used domestically. It reflects the reality that businesses are both building their own tools [as well as] buying off the shelf via closed-model companies like Anthropic and OpenAI.

Context: Confirmed by mid-2026: CNBC reported Western companies increasingly adopting Chinese open-weight models on cost grounds, and OpenRouter token-traffic data showed GLM-5.2 usage climbing faster post-launch than any predecessor — Chinese open models are in production US use at scale.

OpenAI's o3 and o4-mini make reasoning models agentic

The new reasoning models could invoke every ChatGPT tool — web search, Python, visual reasoning, image generation — inside a chain of thought. OpenAI reported the o3 cost-performance frontier strictly improved on o1's, evidence that the reasoning axis was compounding rather than saturating.

Llama is no longer the open standard.

Context: Borne out within a year: by August 2025 MIT Technology Review reported Chinese open models (DeepSeek, Kimi, Qwen) had become more popular than Llama, and by mid-2026 the independent record — UK AISI's evaluations and the Artificial Analysis index — showed every leading open-weight model was Chinese until Thinking Machines' July 2026 entry.

Meta's Llama 4 launch is marred by a benchmark-integrity controversy

Scout and Maverick shipped as open-weight multimodal MoE models, but Meta's headline LMArena Elo of 1417 came from an unreleased experimental chat configuration, not the published weights — drawing independent criticism, and diverging evaluations: Artificial Analysis rated the models among the best non-reasoning releases while other independent testers found them far weaker than claimed.

Google's Gemini 2.5 Pro takes the LMArena lead

Released as a free experimental model, 2.5 Pro debuted at #1 on the LMArena human-preference leaderboard by what Google called a significant margin, with reasoning built in — the beginning of a sustained Google run at the top of crowd rankings that lasted more than six months.

DeepSeek releases R1, an open-weight reasoning model priced at ~2% of o1

R1 shipped under an MIT license with API pricing of $0.55 per million input tokens and $2.19 output — versus $15/$60 for the o1 it claimed parity with on math, code, and reasoning benchmarks — alongside six distilled smaller models. The release triggered a global market repricing of AI economics within a week and made a Chinese lab, for the first time, the reference point in the frontier conversation.

DeepSeek-V3: frontier-adjacent open weights at a fraction of the training cost

A 671B-parameter (37B active) mixture-of-experts model with open weights, trained — per DeepSeek's technical report — in just 2.788 million H800 GPU-hours on 14.8T tokens, with day-one inference support on AMD and Huawei Ascend hardware. The claimed training economics, on export-restricted hardware, previewed the disruption its successor would cause a month later.

OpenAI's o1 opens the reasoning-model era

o1-preview and o1-mini — trained, per OpenAI's research lead, with a new optimization approach that spends more compute thinking before responding — launched at $15/$60 per million tokens, roughly six times GPT-4o's price. Test-time compute became the industry's second scaling axis, and 'reasoning' variants became standard across every major lab within a year.

GPT-4o mini resets the price floor at $0.15 per million tokens

OpenAI's small-model replacement for GPT-3.5 Turbo shipped at $0.15/$0.60 per million tokens — an order of magnitude cheaper than prior frontier models, and by OpenAI's own accounting a 99% cost drop per token versus its 2022-era text-davinci-003. The collapse in the price of a fixed level of capability was now explicit vendor strategy.

Apple Intelligence puts a ~3B-parameter model on the iPhone

Apple's WWDC announcement paired a ~3 billion parameter on-device model (compressed to ~3.7 bits per weight, with swappable per-feature LoRA adapters) with a larger server model on Apple silicon. In Apple's own evaluations the on-device model beat comparable small open models — the clearest early evidence for a second, edge-scale track in the model race.

Anthropic releases the Claude 3 family

Haiku, Sonnet, and Opus, priced at $0.25/$1.25, $3/$15, and $15/$75 per million tokens respectively, all with 200K context. Anthropic's published benchmark table showed Opus beating GPT-4 across common evaluations — the first credible claim on GPT-4's crown since its release a year earlier.

Claude 3 Opus has self-reported benchmark scores that consistently beat GPT-4. This is a really big deal: in the 12+ months since the GPT-4 release no other model has consistently beat it in this way.

Context: Anthropic's published table did show consistent wins over GPT-4's reported scores, and the retrospective independent record supports the assessment: on Artificial Analysis's current Intelligence Index (v4.1), Claude 3 Opus (12, estimated) scores above GPT-4 (7, estimated).

Google's Gemini 1.5 Pro debuts the million-token context window

Announced with a standard 128K context and a 1-million-token window in private preview — the longest of any large-scale foundation model at the time — built on a mixture-of-experts architecture Google said was cheaper to train and serve. Context length became a new axis of frontier competition alongside raw capability.

Mistral releases Mixtral 8x7B, open weights under Apache 2.0

A sparse mixture-of-experts model (46.7B total parameters, 12.9B active per token) under a genuinely permissive license — no research-only gate, no commercial restrictions. Mistral claimed it outperformed Llama 2 70B with 6x faster inference and matched GPT-3.5 on standard benchmarks, making it the strongest truly-open model of its moment and the European entry in the frontier race.

Google announces Gemini, its first natively multimodal model family

Three sizes — Ultra, Pro, and Nano — with Nano shipping on-device on the Pixel 8 Pro, the first frontier-lab model engineered into a phone. Google claimed Ultra exceeded state-of-the-art on 30 of 32 academic benchmarks and was the first model to beat human experts on MMLU at 90.0% — vendor figures that independent same-day coverage noted rested on very close margins.

I think we're substantially ahead on 30 out of 32±

Context: The Verge, reporting the same day, noted the claimed margins over GPT-4 were 'mostly very close' — the 30-of-32 tally is technically Google's own benchmark table, and the practical gap it implied did not match the headline framing.

Meta releases Llama 2 with open weights for commercial use

Weights and starting code for pretrained and fine-tuned versions, free for research and commercial use, with Microsoft as preferred distribution partner via Azure. Meta reported over 100,000 researcher access requests for the first LLaMA. Open-weight frontier-chasing became a deliberate corporate strategy rather than a leak — though the license's restrictions meant 'open source' was contested from day one.

OpenAI releases GPT-4

A large multimodal model accepting image and text input, released via ChatGPT Plus and an API waitlist at $30/$60 per million tokens (8K context; the 32K version cost $60/$120). OpenAI also open-sourced its evaluation framework, OpenAI Evals. GPT-4 became the reference point the rest of the industry measured against for the next year.

Anthropic opens access to Claude

After a closed alpha with partners including Notion and Quora, Anthropic opened Claude via chat interface and API, in two tiers: Claude and the lighter, cheaper Claude Instant — putting a second US lab's frontier-adjacent model into the market the same week as GPT-4.

Meta's LLaMA weights leak onto the open internet

Roughly two weeks after Meta announced LLaMA as a research-access release requiring an application, the model weights appeared as a downloadable torrent on 4chan. Observers split between warning the technology would be misused and arguing open availability accelerates safety research — the first major access rupture of the frontier era, and an accidental preview of the open-weights ecosystem to come.

OpenAI releases ChatGPT as a free research preview

A conversational model fine-tuned with RLHF from a GPT-3.5-series base went live to the public at an effective price of $0. OpenAI's own launch notes disclosed the system 'sometimes writes plausible-sounding but incorrect or nonsensical answers.' The release turned large language models from a developer API into a mass consumer phenomenon and set off the competitive race that defines this page.

DeepMind's Chinchilla finds the era's giant models undertrained

DeepMind published its compute-optimal training analysis: a 70B-parameter model trained on 1.3 trillion tokens outperformed far larger models like the 280B Gopher and 530B Megatron-Turing NLG. The finding — that the era's models were oversized for their compute budgets and undertrained on data — reset every lab's scaling recipe and foreshadowed the industry's turn toward smaller, cheaper-to-serve models.

OpenAI opens the API era with GPT-3

OpenAI released its first commercial product: a private-beta 'text in, text out' API serving the GPT-3 family. It explicitly chose API access over open-sourcing the weights, arguing access can be revoked if misused — and that the models were so large and expensive to run that only big companies could otherwise afford to deploy them. The pattern set here — frontier capability behind a metered API — defined the market for years.