Today's Headlines
- Qwen publishes the weights of its vision-language model Qwen3.8-27B under Apache 2.0, updated exactly at midnight JST on August 15
- Computer History arrives in ChatGPT and Codex, recording desktop actions on macOS and turning them into memory
- Z.ai announces GLM-5.3 on the same base model as GLM-5.2, with a 50% gain on its internal coding benchmark
- Google adds a setting that turns off the visible watermark on its AI generations, while SynthID and C2PA stay attached
- A 155-year-old umeboshi shop builds its own business portal with Claude and NetSuite, without an engineer on staff
Today's three stories gather around where a system gets improved. Three companies aimed at the same goal and put their hands in three different places.
What Qwen changed is distribution. The weights of a 27B vision-language model now sit within reach, under a license that permits commercial use.
What OpenAI changed is the raw material of memory. Actions taken in front of the screen are recorded and turned into something a later conversation can read.
What Z.ai changed is the finishing. The base model stayed where it was, and the company says the gains came from the scale of post-training alone.
Today's Top Three
Qwen publishes the weights of Qwen3.8-27B under Apache 2.0
Qwen published the weights of Qwen3.8-27B, a 27-billion-parameter vision-language model, on Hugging Face under the Apache 2.0 license.
The timing matched the announcement. This column reported on August 13, citing PC Watch, that the model was scheduled to arrive as open weights at midnight on August 15. Checking the Hugging Face API, the editors found the repository last modified at 15:00:01 UTC on August 14 — midnight JST on August 15 — exactly the hour that had been promised.
The specification shows both scale and range. The model carries 27B parameters across 64 layers, and the model card gives the context length as "262,144 natively and extensible up to 1,000,000 tokens" for a native vision-language model that reads images and video.
The benchmark figures come from the model card. SWE-bench Pro reads 61.7 against 53.5 for the previous Qwen3.6-27B and 53.4 for Claude Opus 4.6 Max, and QwenSWEBench reads 79.0 against 49.3 and 63.8 in the same order.
The same table also records the losses. Terminal Bench 2.1 (Terminus) puts Qwen3.8-27B at 73.0 against 78.2 for Opus 4.6 Max, and NL2Repo-Bench puts it at 42.3 against 47.6, with Opus ahead on both.
The license sets the terms of local use. Apache 2.0 permits commercial use, which puts the model on the list for companies that want to run it inside their own environment. A hosted version on Qwen Cloud is listed as coming soon.
- Qwen/Qwen3.8-27B (Hugging Face)
- Qwen3.8-27B weights published, with reports of it beating Opus 4.6 Max on some benchmarks (ITmedia AI+, in Japanese)
Computer History arrives in ChatGPT and Codex, recording desktop actions as memory
OpenAI shipped Computer History, a feature that records what happens on screen and turns it into memory for ChatGPT and Codex.
What gets recorded is the action itself. The documentation lists "clicks, typing, keyboard shortcuts, app switches, and context that macOS exposes through its accessibility system."
The exclusions are stated with equal clarity. Screenshots, screen recordings, microphone input and system audio stay outside the capture, and private-mode browsing stays outside it as well.
The environment is a narrow one. The feature runs in the macOS ChatGPT desktop app for Pro, Business and Enterprise plans, and it ships off by default. On Business and Enterprise, an administrator grants it in workspace settings before an individual can switch it on.
Two further conditions apply. Memories is a prerequisite, and access through an API key or Amazon Bedrock sits outside the feature. Availability currently stops short of the European Economic Area, Switzerland and the United Kingdom.
The use reaches both products. The documentation says summaries of recorded activity can be referenced from ChatGPT and from Codex, letting a user pick up where the work left off. ITmedia reports that repeated work gets detected and offered back as a skill or an automation.
- Computer History (OpenAI Docs)
- ChatGPT records Mac activity and uses it as context (ITmedia AI+, in Japanese)
Z.ai announces GLM-5.3, built on the same base and lifted by post-training
Z.ai announced GLM-5.3 on August 14, a model aimed at coding and security work.
The foundation stayed where it was. The official documentation states that the model "uses the same base model as GLM-5.2 — all improvements come from post-training."
The gains are given in the company's own numbers. Z.ai reports a 50% improvement over GLM-5.2 on Z.ai Code Bench and claims open-source state of the art on Terminal-Bench 3.0 and Agents' Last Exam (CLI). Context length runs to 1M tokens with a maximum output of 128K.
Security benchmarks carry figures as well. The documentation puts GLM-5.3 at 84.5% on CyberGym, slightly ahead of Mythos 5 at 83.8%. The figures for GLM-5.2 at 77.2% on the same benchmark, 54.4% on ExploitBench, 105 tasks within two hours and 130 within six on ExploitGym, and 2,436 vulnerabilities found across 269 real codebases with 1,097 rated high to critical, all come from the PC Watch article.
Access stays limited for now. The model is available to paying GLM Coding Plan users, with the API listed as coming soon. Z.ai puts the weight release within two weeks of the announcement, and the zai-org organization on Hugging Face still lists GLM-5.2 as its newest entry today.
- GLM-5.3 (Z.ai Docs)
- "A match for Fable 5": GLM-5.3 arrives, with big gains from post-training alone (PC Watch, in Japanese)
Read the three together and the place of improvement comes apart. One side hands over the weights, one side gathers what happens at the desk and turns it into context, and one side leaves the base alone and thickens the finish — all on the same day.
Other Developments
Models
Writer introduced Palmyra X6, a new flagship model paired with a rebuilt agent harness. TechCrunch reports that it is a post-trained version of Z.ai's open-source GLM-5.2, and that Writer measured an average 52% cost reduction, 48% speed gain and 10% quality gain across its own agent products. The company's internal research found that harness-side changes cut costs more reliably than the choice of model, by an average of 40%.
Mixedbread announced Toast 1, an agent specialized for search and offered through an API only. It decomposes a query into sub-queries and handles everything from evidence gathering to source inspection. The company reports that GPT-5.6 Sol paired with Toast 1 reached 70% accuracy at about $1.15 per task on the OfficeQA Pro V2 financial analysis benchmark, against 60% at $4 per task for the previous best, Claude Fable 5.
Products
Google will add a setting within days that turns off the visible watermark on generated images, video and music. TechCrunch reports that it covers the Nano Banana, Omni and Lyria models and lives under Settings > Media Watermark. The invisible SynthID identifier and C2PA metadata stay attached regardless of the setting.
Chinriu Honten, an umeboshi specialist founded in 1871 in Odawara, Kanagawa, connected the cloud ERP NetSuite with Anthropic's Claude so that a managing director outside engineering could build inventory management, shipment visibility and automatic product specification sheets in-house, ITmedia reports. The work began in April 2026, and drafting a product specification sheet had previously taken two hours per case. The team built a simple portal that a craftsman can operate in a few clicks, and a dashboard for overseas market strategy through the same route. As background for English-language readers, umeboshi are salted, pickled plums, a staple of Japanese home cooking.
DeepSeek released DeepSeek Harness, an agent platform in which models, tools, skills, sessions and sandboxes all attach as plugins, as a developer preview on GitHub under the MIT license. PC Watch reports that the core Cordis framework handles plugin management alone. The release landed on August 13 China time and remains a preview, with breaking changes possible.
- DeepSeek releases DeepSeek Harness, an all-plugin agent platform, in preview (PC Watch, in Japanese)
Anthropic published a practical guide to getting more out of a Claude Code session. Output tokens cost roughly five times input, and prompt caching charges full price on the first write and 0.1x on later reads, with the cache expiring after an hour, so the guide recommends /compact before a break. It also recommends /clear at task boundaries, @ mentions for file references, and several shorter sessions in place of one long one.
Research
Anthropic announced that future Claude models will embed a watermark in generated text, built on Google DeepMind's SynthID-Text. Verification works through statistical patterns in word choice, and the company says the mark stays invisible to readers while quality and speed hold steady. Only a holder of the dedicated key can run the check, and the move sits alongside regulatory work such as the EU AI Act.
Google published its work on making homomorphic encryption practical, running inference over data that stays encrypted. The open-source compiler stack HEIR sits at the center, letting a trained model operate on encrypted input. The company contrasts the purely cryptographic guarantee with hardware-dependent approaches such as secure enclaves.
Policy
In a Connecticut civil case, a self-represented plaintiff who suspected the court of processing filings with AI hid instructions favorable to himself in white, minuscule type inside two documents, Ars Technica reports. Court staff caught it after noticing that the line and character spacing differed from his earlier filings. The judge stated that the court handles filings without AI, and sanctioned the plaintiff by barring electronic filing and allowing paper submissions only.
A partly redacted Anthropic document titled "Risk Report: August 2026," running 186 pages, surfaced through Hacker News on August 14. It records updates to the Responsible Scaling Policy thresholds covering automated AI research and development and the creation of biological and chemical weapons. The publication date sits outside the PDF itself, so what stands confirmed is its surfacing on August 14.
Business
Indonesia's Ministry of Communication and Digital Affairs, Indosat, NVIDIA and Gadjah Mada University opened the country's first university-based AI technology center in Yogyakarta. It gives researchers and students enterprise-grade GPU compute for work on tuberculosis, agriculture and disaster preparedness. NVIDIA's blog cites more than a million new tuberculosis cases a year in the country and roughly 30% of the workforce in agriculture as the background for those research themes.
An analysis from the energy research firm Noreva suggests that U.S. hyperscalers leaning on natural gas for AI data center power could regret it if prices climb, TechCrunch reports. The analysis sees room for gas in some regions to move from $2–4.50 per MMBtu to above $10. Meta is planning 7.5 gigawatts of gas generation in Louisiana and Amazon 7.6 gigawatts in Texas, with fuel accounting for about half the cost of the power.
Apple is building its own generative AI model for the Chinese market with technical help from Alibaba, The Verge reports, citing people familiar with the work. China's Cyberspace Administration registered Apple's generative AI service in July, which could make Apple the first U.S. company to offer its own model there. The model is expected to sit alongside third-party models such as Alibaba's Qwen inside the Chinese version of Apple Intelligence, with a launch expected within months.
OpenAI and Anthropic have both cut prices, Ars Technica reports. The outlet says OpenAI dropped input pricing for GPT-5.6 Luna from $1 to $0.20 per million tokens and trimmed another model, Terra, by about 20%. Anthropic's Claude Opus 5 runs at $5 input and $25 output per million tokens, a figure the editors confirmed on the company's official pricing page, which is half the $10 and $50 of Fable 5. The outlet points to pricing from Chinese rivals such as DeepSeek and Moonshot AI as the pressure behind the moves.
Other
Nankai Electric Railway and Hitachi began building a system that generates crew and rolling stock operating plans automatically, using Hitachi's CMOS annealing, a technique that emulates a quantum computer. Testing during fiscal 2025 confirmed that crew planning, which had taken several months, comes down to about a week, and rolling stock planning from about 20 days to a few days. Coverage expands from the Nankai and Airport lines to the Koya and Semboku lines, with service targeted during fiscal 2027.
Source: Selected by the editors from the AI news inbox (22 items collected on August 15, 2026 — 8 primary, 14 secondary).