Today's Headlines
- NVIDIA's own AVO harness scores 100.00 RHAE on the public set of ARC-AGI-3 — Claude Opus 5 alone sits at about 30%
- Anthropic opens Claude Mythos 5 for enterprise vulnerability scanning and announces a fund of $35 million in Claude credits
- Excel's =COPILOT function will be withdrawn on September 14, 2026 without reaching general availability
- Orbital data-center company Starcloud raises another $250 million at a $2.3 billion valuation
- MUFG Bank says AI flowcharts should cut per-case standardization work from 11 person-days to one
Today's three stories all turn on what sits around a model rather than inside it. The wrapping takes a different form in each one, and each form decides something different.
On the benchmark, the wrapping decided the score. What NVIDIA measured was a system with memory and supervision attached, rather than a model on its own.
On dangerous capability, the handover drew the line. What Anthropic opened is a counter for collecting defensive output, rather than a door to the model itself.
In the spreadsheet, the withdrawal set a precedent. What Microsoft is pulling is the function as an entry point, while the AI itself stays in the side pane.
Today's Top Three
NVIDIA's AVO harness clears the public set of ARC-AGI-3
NVIDIA has reported that its own agent harness, AVO, cleared the public set of ARC-AGI-3 with a perfect score.
The report is a post dated August 21 on the NVIDIA Technical Blog. By the company's own measurement, AVO completed all 183 levels across 25 environments in the benchmark's public set, for an RHAE score of 100.00.
The comparison starts with the model on its own. Run directly against the same public set, Claude Opus 5 scored roughly 30%.
AVO is a wrapper rather than a replacement. It adds persistent memory and a supervisor agent that oversees subordinate agents, holding state across work that runs for a long time.
There is an efficiency figure as well. NVIDIA says AVO finished the same public set using 12% fewer environment actions than a comparison system called VISTA.
The company also ran it outside the benchmark. In a GPU kernel optimization experiment, AVO operated continuously for seven days, explored more than 500 optimization directions and committed 40 kernel versions.
NVIDIA measured the output of that run too. The multi-head attention kernel it produced ran up to 3.5% faster than cuDNN and up to 10.5% faster than FlashAttention-4 on an NVIDIA DGX B200, by the company's own benchmarks.
The provenance of these numbers is worth holding onto. Every figure above is NVIDIA's own measurement on the public set, and independent replication is still to come.
Adel El Hallack of NVIDIA describes AVO as something other than a new product. By his account, AVO is part of the open set of harness-building components the company publishes under the Nemo brand.
- NVIDIA AVO Reaches 100% on ARC-AGI-3, Demonstrating a Frontier-Level General-Purpose Architecture for Long-Horizon Autonomous Agents (NVIDIA Technical Blog)
- Nvidia just showed that the harness, not the AI model, is now the real hero (TechCrunch)
Anthropic opens Claude Mythos 5 to defenders and announces a $35 million credit fund
Anthropic has announced that its top model, Claude Mythos 5, is now available on the defensive side of cybersecurity.
The announcement is a post dated August 21 on the claude.com blog. What became available the same day is vulnerability scanning inside Claude Security for enterprise customers, offered as a public beta for Claude Enterprise.
The pricing stays inside the existing frame. There is no surcharge beyond ordinary token billing, results come back as a CWE classification, a confidence level, a severity rating and a suggested fix, and a human has to review and approve before anything is applied.
The design point is the partition between the user and the model. Users collect the defensive artifact — a patch, an alert — instead of reaching the model itself.
The fund is made of something other than cash. The $35 million behind the new Defender Advantage Fund is Claude credits, going to organizations that help open-source maintainers secure their code.
How it will be distributed is still being settled. Anthropic writes that it will start with a small number of large pilot grants and name the recipients "in the coming weeks."
Parts of the announcement describe work still ahead. Building Mythos 5 into partners' cyber defense products is at the "working on it" stage, and the expansion of the Cyber Verification Program lies further out.
The company has set an order for that expansion. It plans to widen the program to the broader dual-use capabilities of Opus and Sonnet within weeks, with the Mythos class following after that.
There is a separate line of direct support. Under Project Glasswing, which began in April, Anthropic has made $4 million in direct donations to open-source security organizations.
This move sits opposite a story this briefing carried the day before. The ciphertext route found in Grok has been waiting on a fix since Adversa AI reported it on June 3, and the firm says it could still reproduce the chain on August 19.
The same class of capability is being handed out through a narrow counter in one place and left running inside a shipped product in the other. Whether the handover gets designed at all is what changes where the capability points.
- Bringing the cybersecurity capabilities of Claude Mythos 5 to more defenders (Anthropic)
- Cryptographic Context Injection: Grok Data Theft (Adversa AI)
Excel's =COPILOT function is withdrawn on September 14 without reaching general availability
Microsoft is retiring Excel's natural-language AI function, =COPILOT, on September 14, 2026.
The wording on the company's support page is that it "will no longer be available." The function let users write an instruction in a cell in plain language and get a result back, and it stayed inside the Frontier and Microsoft 365 Insider programs to the end, folding before it reached general availability.
Access was narrow while it lasted. Work and school accounts needed a Copilot add-on license, and among personal accounts only Microsoft 365 Premium subscribers could use it.
The calculation itself carried limits. Recalculation was capped at 100 per ten minutes, and workbooks labelled Confidential or Highly Confidential fell outside the function's reach.
This is the part that invites a misreading. Rather than pulling AI out of Excel, Microsoft is pointing users to the generally available Copilot in Excel side pane.
The start date comes from a publication rather than the vendor. PC Watch dates the Insider preview to August 2025; Microsoft's support page gives no start date at all.
Seen from an enterprise AI rollout plan, this one stands as a precedent. A capability that had reached the level of the spreadsheet's own formula language was pulled before it reached the general-availability stage.
- Microsoft shelves Excel's
=COPILOTAI function before general rollout (PC Watch) - COPILOT function (Microsoft Support)
Other Developments
Models and APIs
Ornith released Ornith-1.5 under an MIT license, a line built from Qwen3.5 and Gemma 4 and extended with a self-improvement loop. The flagship 397B model runs close to Claude Opus 4.8 on coding and agentic benchmarks, and the split is one win each: Terminal-Bench 2.1 goes to Ornith at 86.1 against 85.0, while DeepSWE goes to Opus 4.8 at 59 against 56. The weights are already on Hugging Face and carry no gate.
Unsloth announced its Dynamic v3.0 quantization technique and published several compressed builds of Qwen3.8-27B. The 6.2GB figure belongs to the smallest 1-bit build alone (UD-IQ1_S), which keeps roughly 72% of top-1% accuracy and cuts 89% of the original size while accuracy on 32-token predictions falls from about 25% to under 8–10%, with failed tool calls and looping reported. The 2-bit build (UD-Q2_K_XL, about 9.83GB) holds accuracy better and handles work such as HTML generation steadily.
Products
OpenAI released a plug-in that connects the macOS ChatGPT desktop app to Apple's Messages. It can search and sort a message inbox and draft or send replies, it runs only on Apple Silicon builds, and the only chats that support it are Codex and ChatGPT Work. Choosing "Always allow sending to this chat" removes the final confirmation before a message goes out, and OpenAI itself urges caution about that setting.
Anthropic launched Claude Academy, a free learning site with courses on Claude and on generative AI in general. The company says it built the material from the 4D AI Fluency Framework and the agent-management practices it uses in its own onboarding, and that it weighted the courses toward judgement — which work to hand to AI and which to keep — over step-by-step technique. Translation is in alpha, and Japanese is among the languages covered.
Digital Commerce, the operator of the Japanese adult marketplace FANZA, plans to launch FANZA Studio, a platform covering the production and release of AI-made adult content, with an early beta preview starting August 24. Copyright in generated work, age verification and platform liability all land on the same service at once, so the way it is run will set a reference point.
Research
Google DeepMind entered a research partnership with Fenris Creations, the studio behind EVE Online, an MMO that has run for more than twenty years. The aim is to point a persistent virtual world at the problems current models handle poorly: continual learning, long-term memory, long-horizon planning and complex multi-agent economies. The work starts in offline environments separated from live players and in EVE Frontier, with integration into the main game considered only once the capability has matured.
Nari Labs published an open-source implementation and benchmark for the speech model Qwen3-TTS (1.7B, CustomVoice), running it at 10 requests per second on a single NVIDIA H100 SXM with first-audio latency under 50 milliseconds at p95. The company puts the cost at about $2 per million characters, well below commercial APIs. Every figure comes from Nari Labs' own measurements, and third-party verification is still to come.
Policy and courts
In EPAM Systems v. Rao, a trade-secret case pending in the Eastern District of Pennsylvania, the treatment of generative AI output itself has surfaced as an issue. For the documents EPAM wants sealed, the respondent states in a sworn declaration that they were made by entering prompts into a personal Gemini account (on a premium plan, expensed to the company) and exporting the result to Google Docs. EPAM's witness contends instead that they were produced by feeding confidential information to generative AI, and the two accounts stand directly opposed.
Eight newspaper publishers, including the New York Daily News and the Chicago Tribune, filed a first amended complaint in their copyright suit against Microsoft and OpenAI entities. OAI Corporation, LLC and OpenAI Holdings, LLC — entities created in the restructuring — join as defendants, and five claims survive: direct infringement, vicarious infringement, removal of DMCA copyright management information, and trademark dilution under the Lanham Act and state law. Claims for contributory infringement and common-law unfair competition were dismissed by earlier court order.
OpenAI began a preview of Private Safety Processing, a safety mechanism built on top of Zero Data Retention, for a subset of customers. Where the existing arrangement evaluates risk one exchange at a time, the new one detects risk automatically across a whole set of related exchanges, and OpenAI says it receives only minimal signals while leaving the content itself untouched. A wider rollout is planned for September.
As Meta's AI glasses spread, unofficial apps that detect people wearing them keep appearing, and schools, courthouses and restaurants are turning them away, Ars Technica reports. The Electronic Frontier Foundation objects that the recording indicator light can be covered with a sticker, and warns about the possibility of Meta adding deeper features such as face recognition.
Business
Starcloud, which builds satellites that run AI inference in orbit, added a $250 million extension to its March Series A of $170 million, reaching a $2.3 billion valuation. Manhattan West led the round, with NVIDIA, Cisco, Benchmark, EQT and Soma among the participants. The money goes to a large manufacturing site and to Starcloud-3, its biggest craft yet, slated to fly on SpaceX's Starship — and the tightening supply of launch slots, with Falcon 9 due to retire in 2028, shapes what the money is for.
MUFG Bank told AWS Summit Japan 2026 that AI flowcharts reconciling the clerical procedures that differ across its branches in 30 countries should cut the work per case from 11 person-days to one. The 90% figure is a projection rather than a measurement taken after deployment, and the one to two person-months spent consulting stakeholders stays where it is. Attempts through May 2025 using a general-purpose LLM alone failed, producing flowcharts at different levels of detail per branch that could not be compared; teaching the business knowledge first, through an ontology and knowledge graph built from roughly 20,000 triples, is what changed the outcome.
The US Department of Justice has spent about a year investigating Andreessen Horowitz over partners holding board seats at portfolio companies that became competitors, according to reports. The companies are Databricks and Fivetran, and the provision under consideration is a 112-year-old antitrust clause with almost no history of application to venture capital.
Card and payment data from the expense platform Ramp puts the share of US companies paying for an Anthropic subscription or tokens at 43.5% as of July 2026, against 39.7% for OpenAI. Set against 41% and 39.5% in May, OpenAI has closed some of the gap — though the figures rest on a sample of more than 70,000 companies skewed toward tech, and Ramp changed its methodology between April and May 2026, so comparisons across that change need care.
Source: selected by the editors from the AI news inbox (45 items collected on August 22, 2026 — 9 primary, 36 secondary).