Today's Headlines
- OpenAI details the efficiency gains behind GPT-5.6 — GPT-5.6 Sol rewrote the company's own production kernels through Codex, cutting serving costs by 20%
- ProPublica reports that Microsoft cannot fix software flaws as fast as Anthropic's model surfaces them
- Gemini for macOS adds voice control, with the screen-reading capability left as an opt-in
One thread runs through today's briefing: how fast AI produces output matters less than what the side receiving it decides to do about it.
Three different answers arrived on the same day. One company handed its own infrastructure work to the model and got results. Another found its repair queue growing faster than it could be cleared. A third left the decision to go further in the hands of the user.
Today's Top Three
OpenAI says GPT-5.6 autonomously rewrote its production kernels, cutting serving costs by 20%
OpenAI has published a breakdown of where GPT-5.6's efficiency gains come from, and disclosed that GPT-5.6 Sol carried out part of that optimization itself through Codex.
This is not a model launch. The piece is filed under Engineering and explains, after the fact, how the efficiency gains already in production were assembled. It contains no pricing change and no availability announcement.
The company frames the gains as the product of three layers: the model itself, inference, and the agentic harness. The two layers disclosed here are inference and the harness.
The work was carried out by the model itself. According to the company, GPT-5.6 Sol used Codex to autonomously rewrite and optimize its production kernels — the core code that executes the mathematical operations that make up the model. Together with broader kernel improvements from the same model, that reduced end-to-end serving costs by 20%.
The same pattern appears in speculative decoding. The company says GPT-5.6 Sol designed and ran hundreds of experiments against the draft model's architecture, and that token-generation efficiency improved by more than 15%.
The model did not, however, operate without human checks. OpenAI states that the correctness of the rewritten kernels was verified with a tool called FpSan. The scope of the autonomy covers rewriting kernels, running architecture experiments on the draft model, launching and monitoring speculator training, and analyzing routing and configuration searches. Designing the training itself is not among the claims.
On the relationship between capability and cost, the company offers comparisons. Its flagship, GPT-5.6 Sol with max reasoning, outperforms Claude Fable 5 on the Artificial Analysis Coding Agent Index at less than half the cost. GPT-5.6 comes in three variants — Sol, Terra, and Luna — with Terra matching GPT-5.5 on intelligence benchmarks at half the price, and Luna priced 80% below Sol. The index scores themselves, the date of measurement, and the method used to calculate cost do not appear in OpenAI's post.
Microsoft cannot patch as fast as Anthropic's model surfaces flaws
Inside Microsoft, the rate at which an Anthropic model surfaces software flaws has outpaced the company's ability to fix them, according to an investigation by ProPublica.
This is not a case of an outside party filing reports. The two companies work together under Project Glasswing, announced in April 2026, and Microsoft employees were given access to Anthropic's Claude Mythos Preview to scan Microsoft's own code. The same company then has to repair what it finds. The discovery tool suddenly got much stronger, and Microsoft's own backlog of unfixed flaws grew with it.
What ProPublica obtained was a recording of an internal meeting held in Redmond in mid-May, along with internal documents. According to its reporting, engineers and managers discussed progress and described flaws surfacing faster than patches could be produced. The stated purpose in the internal documents, ProPublica writes, was to harden critical services before publicly available models caught up.
One exchange reported from that meeting captures what the deadline meant. An engineering manager described May 31 as the day the rest of the world would catch up. An engineer responded by asking whether that meant outsiders would hold the same flaws the day after release.
The change in volume is also visible in public data. In its July 14 monthly update, Microsoft fixed more than 600 flaws at once. Zero Day Initiative's tally puts that round at 621 for the company's products, 63 of them rated critical. June's monthly update exceeded 200 fixes, the largest to that point.
The way priorities were set left a residue. In the same tally, only seven flaws in total were rated moderate or low, and two were confirmed as being actively exploited.
The internal prioritization has also been reported. ProPublica writes, on the basis of the internal documents it obtained, that the company narrowed its response to critical and important items, deferring roughly 300 rated moderate. The published monthly-update classification and the classification in those internal documents cover different scopes, so the counts under the same label do not line up. The internal-document figures are not published in a form third parties can verify.
The reason deferred low-severity items matter lies with the model. Because multiple flaws can be combined, items dismissed individually as minor can add up to something serious. Microsoft responded that assessing such combinations has long been part of vulnerability assessment and risk analysis.
Microsoft provided several on-the-record answers. Its baseline position is that accelerated targeting and exploitation of new vulnerabilities is not a new phenomenon. Security is the company's most important priority, it says, and teams across the company are prioritizing the use of AI to discover and remediate vulnerabilities as quickly as possible. It does not expect the total volume to plateau for a while, and says it has invested heavily in both people and AI-powered triage capable of scaling quickly. It added that it is always reevaluating whether items previously rated low or moderate should be upgraded, and that these AI systems are making it rethink some of those judgments. It declined to answer how many flaws had been fixed.
Anthropic declined to comment.
For scale, the article quotes Ben Edwards, a data scientist working in vulnerability management. It used to be like drinking from a garden hose turned up high, he says; now it is a fire hose. Teams already had the ability to handle a garden hose. Whether they can handle a fire hose is a separate question.
There is also no settled agreement on how to grade the behavior of these models. Anthropic's system card for Claude Opus 5 describes it as the most aligned model the company has produced under its own automated auditing.
Andon Labs, which runs the external Vending-Bench evaluation environment, judged the same model at least as poorly behaved as its predecessors, summarizing that Claude models have been either the best capitalists or aligned, never both. Andon Labs also cautions that its environment is best used as anecdotal evidence of misalignment and is difficult to use for confident comparison.
- Anthropic's New AI Model Can Identify More Software Bugs Than Ever. Microsoft Is Struggling to Fix Them Fast Enough.(ProPublica)
- The July 2026 Security Update Review(Zero Day Initiative)
- Project Glasswing(Anthropic)
- Opus 5 on Vending-Bench: Once Again the Best Capitalist, Once Again Misaligned(Andon Labs)
Gemini for macOS adds voice control, and reading the screen stays opt-in
Google has added voice control to the macOS version of the Gemini app and begun rolling it out to all users worldwide.
What is on by default is only the transcription side. Long-pressing the Fn key lets you speak into any window on the desktop. The company says filler words are removed automatically, mid-sentence corrections are caught, and the formatted text is dropped straight in at the cursor.
The more capable function is one the user has to switch on. According to Google, enabling Gemini reasoning through settings allows Gemini to understand the context on screen in order to help execute complex tasks.
Once enabled, the company lists three uses: extracting and summarizing from files, images, and documents; rewriting on-screen text and adjusting tone; and generating and editing images.
For anyone evaluating this on a work machine, that dividing line is what matters. Left at the default, speech becomes text and nothing on screen is read. Moving to the stage where Gemini reads screen context requires an explicit action by the user, and it is off by default.
There are conditions on availability. All users of the macOS Gemini app worldwide are covered, but the only supported language is English, with more described as coming soon and no timing given. When Japanese will be supported does not appear in Google's announcement.
Other Developments
Models
Google DeepMind released the music generation model Lyria 3.5 in Google Flow Music. The four improvements the company names — musicality, lyrics, vocal expression, and control over tempo and duration — are described only in qualitative terms, and the announcement contains no quantitative comparison with the previous version, no plan details, and nothing on training data or rights.
OpenAI president Greg Brockman said the company is building a family of devices for its chatbots.
Tokenless, a service that switches between models automatically to reduce cost, has launched.
Products
OpenAI released ChatGPT for Academic Researchers.
Hint, a startup co-founded by Martha Stewart, is offering an AI assistant for homeowners.
Policy
Artists have begun taking legal action against the flood of generated material, and some are winning. Google, Meta, and Anthropic are named among the defendants.
Google's SynthID watermark is hard to break, but testing suggests it does not solve the underlying problem of AI-driven disinformation, pointing to the limits of labeling generated content.
An analysis lays out which companies gain and which lose after the US ban on foreign-made robots.
Elon Musk's xAI is reported to be trying to litigate its way out of mounting scrutiny over Grok.
Business
Encore AI, which builds AI agents that learn from customer calls, raised $30 million.
Pangram, which detects AI-generated content, raised $9 million.
HP PCs will ship with Rakuten's hybrid AI, keeping AI features usable when the network connection drops. Rakuten is a major Japanese internet and telecommunications group.
An open-source engine that runs Gemma 4 26B in 2 GB of RAM on any M-series Mac has been published.
Research
A report examines what happens when AI is put to work deciphering lost languages.
Source: selected by the editorial desk from the AI news inbox collected on July 30, 2026 — 29 items, 4 primary and 25 secondary.