Today's Headlines

  • Qwen opens the weights of a Max-class model for the first time — 2.4 trillion parameters total, 95 billion active
  • SpaceXAI releases Grok 4.6 — built around long-running agents, priced at $2 in and $6 out per million tokens
  • OpenAI reports that top-decile firms generate 8.3x the output tokens per user — up from 2.6x in January
  • The Gemini app passes one billion monthly active users — Google's fourteenth product to reach that mark
  • Mercari rolls out Claude Code across the company — permissions split automatically through MDM and IDP

Today's three stories set the reach for higher performance beside the question of how far it has spread. The first two are about how high the ceiling goes; the third is about how far the capability has actually travelled.

What the Qwen team moved is where the top tier sits. A model of a size that had lived behind a commercial API is now out in the open, weights and all.

What SpaceXAI moved is the centre of gravity of a model. The focus has shifted from a single reply to an agent that keeps running for hours.

What OpenAI put out is the story of what happens after distribution. Among companies that can all reach the same models, the way they use them has already pulled apart, and the company shows it with its own data.

Today's Top Three

Qwen opens the weights of a Max-class model for the first time

The Qwen team has published Qwen3.8-2.4T-A95B, a model with 2.4 trillion total parameters and 95 billion active parameters, as open weights on Hugging Face.

This story has a run-up. This briefing tracked the release on August 10 and August 11 as "scheduled for August 12," and on August 12 reported plainly that nothing had appeared on the day itself. The repository on the official Qwen account on Hugging Face was last updated at 10:24 UTC on August 12, which puts the release inside the week the company had signalled.

The architecture is a mixture of experts. The model card describes 92 layers and 512 experts in total, of which 10 routed experts and one shared expert are activated per token. Against a total of 2.4 trillion parameters, 95 billion do the work.

The context length is published as well. The model card gives 262,144 tokens natively, extensible to a maximum of 1,010,000.

On positioning, the accurate thing is to quote the model card directly. It states that "for the first time, Qwen3.8 brings a Qwen-Max-class model to open release." That the top Max tier has been opened weights and all is a claim made by the publisher.

The benchmark figures are all the company's own numbers, as printed on the model card. Terminal Bench 2.1 comes in at 86.6 against 84.6 for Opus 4.8, PaperBench at 93.0 against 90.5 for GPT-5.6 Sol, and IFBench at 82.8, up from 79.1 for the previous Qwen3.7-Max.

The lead holds in some columns and stops in others. In the same table, SWE-bench Pro is 67.7, below the 80.0 recorded for Fable 5. Every figure so far comes from the company's own testing.

Anyone planning commercial use should read the licence first. The licence is named qwen3.8-max, and by PC Watch's account it is free at base, while services with more than 100 million monthly active users or more than $20 million in monthly revenue must display the model name in their interface, and companies with more than $50 million in consolidated revenue over the past twelve months need a separate licence agreement to offer it commercially through a MaaS or similar arrangement. Purely internal use falls outside that requirement.

SpaceXAI releases Grok 4.6, built around long-running agents

SpaceXAI, Elon Musk's AI company, has released a new model, Grok 4.6.

The centre of gravity the company names is work that runs for a long time. Its announcement says Grok 4.6 builds on Grok 4.5 "with a particular focus on long-running agents and more ambitious interactive and visual work."

Its place on the composite index is given in the company's own chart. The announcement puts Grok 4.6 at 61 on the Artificial Analysis Intelligence Index, the same score as GPT-5.6 Sol Max. The same chart shows Fable 5 Max at 62 and the previous Grok 4.5 High at 56.

Pricing is unchanged. The announcement lists $2 per million input tokens and $6 per million output tokens, with the faster variant at twice that.

Availability started immediately. The company says the model is in Cursor and Grok Build from day one, and also reachable through the API and partners including OpenRouter, Vercel and Cloudflare.

The price comparison needs its source kept separate. The line that Grok 4.6 sits "60%+ below Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30)" comes from Artificial Analysis, an independent measurement firm; the company's own announcement stops at its pricing.

Artificial Analysis publishes its own measurements too. It puts Grok 4.6 at a GDPval-AA v2 Elo of 1,753, behind only Claude Opus 5, and reports that the model completes a task in roughly 53 turns and about 0.5 billion input tokens on average, against roughly 103 turns and about 2.0 billion for Claude Opus 5. The context window stays at 500,000 tokens, unchanged from Grok 4.5.

On the same index, that 61 places Grok 4.6 behind Claude Opus 5 (max) at 63 and Fable 5 (max) at 62. The number is the same in both accounts; SpaceXAI writes it from the side of drawing level with GPT-5.6 Sol, and Artificial Analysis from the side of sitting behind the two in front.

OpenAI reports that the gap in enterprise AI use is widening

OpenAI has published two reports arguing that the gap between companies in how they use AI is widening.

The numbers come from the company itself. One is Enterprise Signals, a view of usage across its own enterprise customer base; the other is a companion working paper, How Organizations Use AI: Evidence from ChatGPT. Both rest on data OpenAI holds. Neither is third-party research.

The size of the gap is measured in output per person. The company ranks enterprise customers each month by output tokens per active user, calls the top 10% "frontier firms" and the 45th to 55th percentiles "typical firms." As of June, the former generated 8.3 times the output of the latter, a threefold widening from the 2.6 times gap in January.

The first component is a change in the shape of the work. As of June, Codex accounted for 64% of the combined Codex and ChatGPT output tokens among enterprise customers, which the company reads as a shift from asking questions towards delegating tasks.

The second is a difference in how far the tooling is used. Among weekly active users, 21% at frontier firms use Plugins against 9% at typical firms, and for skills the split is 19% against 3%. The company cites its own internal weekly Plugin usage of 95% as a sign of how much room is left.

The third is which functions are picking it up. Since February, weekly active enterprise Codex users have grown 108x in legal, 41x in sales, 41x in recruiting and 26x in marketing, against 5x in engineering. Functions that started from a small base will naturally show larger multiples, and that should be discounted; even so, the centre of the growth has clearly moved outside software engineering.

The fourth is the distribution inside a company. Six months after adoption, early-career employees were sending 13 more messages a week than executives. The company notes that surveys tend to show the opposite, with leaders reporting the heavier use.

That 108x in legal reaches readers here directly. Contract work, research and the first drafts of formal documents all handle material that cannot leave the building, which is exactly what slows adoption down. If that is the function growing fastest, as the company reports, then whether internal policy and access design have kept up shows up as the gap itself.

Other Developments

Models

Qwen3.8-27B, the small model in the same family, is scheduled to be released as open weights at midnight on August 15, according to PC Watch. It succeeds Qwen3.6-27B and is said to carry over, in lighter form, the coding and agent techniques developed for the larger Qwen3.8 Max. The publication reports that it is sized to run on hardware within reach — a high-end consumer GPU or a Ryzen AI Max PC — and that it is natively multimodal and multilingual.

NVIDIA has released Nemotron 3.5 Lightning, a 30-billion-parameter open model designed for always-on agents. It uses a mixture-of-experts design that keeps 3 billion parameters active at inference, and the company says it scores level with OpenAI's gpt-oss-120b, a model roughly four times its size. It is placed at 24 points on the Artificial Analysis Intelligence Index, 9 points above the previous Nemotron 3 Nano, with output running at about 670 tokens per second. The training data and recipe are released alongside the weights under the OpenMDW-1.1 licence.

Products

Google used its annual Made by Google event on August 12 to announce the Pixel 11 line, the Pixel Watch 5 and a tracking tag called Pixel Tag, along with new Gemini features. The Pixel 11 starts at $899, a $100 increase on the previous model, with base storage doubled to 256GB. Pixel Tag is $29, or $99 for a four-pack, and the Pixel Watch 5 starts at $399.

Google chief executive Sundar Pichai said on August 11 that the Gemini app had passed one billion monthly active users, making it the fourteenth Google product to reach that mark. Usage figures were published with it: 63% of users work with Gemini by voice, 38% of study-related requests include an attached file, and active users on iOS number more than 100 million.

SpaceXAI has opened a beta of Grok Bot, an agent service designed to act as an independent "AI teammate." According to The Verge, the bots can message one another and divide work in a group chat, and a user has to grant a bot sign-in access to their own accounts. It runs on desktop and iOS for paying subscribers including SuperGrok Heavy, with team and enterprise access handled through a waitlist.

Research

Google DeepMind has published SL2T, a multilingual model that converts sign language into text, and has begun shipping it as American Sign Language to English translation in Gboard and Live Transcribe on the Pixel 11. It was trained on more than 50 sign languages and over 100,000 hours of data, and for privacy the design sends only extracted skeletal coordinates to the server rather than the camera feed itself.

Policy

Judge Mehalchick of the US District Court for the Middle District of Pennsylvania issued an order on August 12 requiring, in every case she handles, a "generative AI certification" from any party that used generative AI to prepare a filing. The disclosure covers three points: which AI tool was used, which portions it produced, and whether a human verified their accuracy. It applies to pro se litigants as well as to counsel, and the order states that a violation may draw sanctions.

Business

Mercari, the Japanese marketplace operator, rolled out Claude Code and Claude Cowork across the company in May 2026 and has now described the machinery behind it. It syncs Jamf for device management with Okta for identity, so that settings are issued automatically according to whether the person holding the device is a developer. Non-engineers are given Notion AI and Claude Cowork as a rule, and the company says that "rolling out Claude Code as-is to non-engineers carries very significant risk from a security standpoint." All model access runs through the open-source LiteLLM proxy, which replaces the providers' own non-expiring API keys with short-lived ones.

Manus, the company behind the Chinese AI agent of the same name, announced on August 11 that it is separating from Meta and resuming operations as an independent company. Meta was reported to have acquired Manus for around $2 billion in December 2025, and Reuters reported that China's National Development and Reform Commission refused to approve the deal in April 2026 and ordered it unwound. As part of the separation, data generated by some users on or after December 29, 2025, the period under Meta, will be deleted between August 23 and 24, Singapore time.

Gartner has forecast that up to $234 billion of enterprise software spending will be exposed to agentic AI by 2030, which it puts at roughly 20% of enterprise SaaS spending in that year. The firm's argument is that once agents complete tasks across several systems at once, the premise of per-seat pricing — that revenue rises with the number of users — no longer holds, and that incumbent vendors will need to move from value based on the interface to value based on the outcome.

Source: selected by the editorial desk from the AI news inbox (37 items collected on August 13, 2026 — 16 primary, 21 secondary).