Today's Headlines

  • Alibaba unveils the 2.4-trillion-parameter Qwen3.8-Max — the first open-weight release at Max scale is promised for next week
  • MiniMax publishes the weights of its video model MiniMax-H3 — the license excludes the EU, the UK, South Korea and the US from its territory
  • UNAM, Mexico's largest university, will make roughly 58,000 applicants retake an entrance exam in person after an AI-proctored remote test
  • Tokyo police arrest a man and refer a high school student to prosecutors over sexually explicit deepfakes made from a real woman's photo
  • The EU AI Act's transparency obligations take effect, requiring machine-readable labels on AI-generated and AI-altered content

Two large models moved on open weights in the same day. One promised its maker's first release at top tier for next week; the other, which generates video and audio together, published its weights today. Both arrive as something you can download.

What "open" means, though, differs. One release is still a promise with a date attached. The other is already published, and comes with a contract that draws a line around the countries where it may be used. Being able to obtain a model and being permitted to use it are settled separately.

The third story is not about a model at all. It is about an institution that brought AI in to supervise, and found that the supervision cost it the credibility of the selection itself.

Today's Top Three

Alibaba unveils Qwen3.8-Max at 2.4 trillion parameters, with the first Max-class open weights to follow

Alibaba's Qwen team has formally announced Qwen3.8-Max, a large model with 2.4 trillion parameters.

The model totals 2.4 trillion parameters and is built on the Qwen3.5 architecture. On the company's own coding benchmark figures, its score on FrontierSWE rose from 40.7 for the preceding Qwen3.7-Max to 73.5.

Comparisons against rival models cut both ways. ITmedia reports that Alibaba claims wins over Fable 5 on Terminal Bench 2.1 and over GPT-5.6 Sol on SWE-bench Pro, while trailing both on other benchmarks including DeepSWE 1.1.

API pricing is $2 per million input tokens and $6 per million output tokens. In a demonstration, the company says the model sustained autonomous coding work for more than ten days.

The weights are due the week after the announcement. It would be the company's first open-weight release of a Max-class model, and a mid-size Qwen3.8-27B is set to be published at the same time.

For now, then, what is actually open is the API. The ability to run the model inside your own environment arrives later. What is new here is that the pattern of Chinese labs publishing weights has reached the top of a flagship line.

MiniMax publishes MiniMax-H3's weights, with the EU, the UK, South Korea and the US outside the licensed territory

China's MiniMax has published the weights of MiniMax-H3, a model that generates video together with stereo audio, on Hugging Face.

The model produces video of 4 to 15 seconds at up to 2K and 24 FPS, with 32 kHz stereo audio generated alongside it. Aspect ratios run from 21:9 to 9:16.

At its core is H3-Omni-Transformer, a 33-billion-parameter dense single-stream Transformer that uses the pretrained weights of Qwen3-VL-32B as its text encoder. Inference runs on the major frameworks, including Diffusers, ComfyUI, SGLang and vLLM.

The point that matters in practice is the license rather than the benchmark. The accompanying MiniMax H3 Community License Agreement carries a license date of August 2, 2026, and states expressly that its scope is limited to what it calls the Applicable Territory.

That territory is defined as worldwide excluding the Excluded Territories, and the Excluded Territories are named as the European Union, the United Kingdom, the Republic of Korea and the United States of America. Japan is not among them, and so falls inside the licensed territory.

The restriction is not confined to distribution. The agreement states that use, reproduction, modification, distribution and display of the works, or of their outputs, outside the Applicable Territory is not authorized under it.

Commercial use carries two further conditions. A licensee whose commercial products and services generate more than $20 million in yearly revenue must obtain separate prior written authorization from MiniMax, and any commercial product or service built on the model must display "MiniMax H3" prominently in its user interface.

The excluded regions are not shut permanently. The agreement says MiniMax will continuously evaluate the applicable laws, regulations and compliance requirements for those territories, and invites anyone there who wants to deploy the models to make contact about a license granted on the basis of robust controls and guardrails.

Governing law is set out as well. The agreement is governed by the laws of the Hong Kong Special Administrative Region of the People's Republic of China, with disputes subject to the exclusive jurisdiction of that region's courts.

The same weights, in short, are treated differently depending on where they are used. This is written as a private contractual term, separate from any public-law regulation a given country may impose. For a company weighing in-house video generation, the questions are not only where it is incorporated, but also how large its revenue is and in which territories the output will be used.

UNAM's AI-proctored remote entrance exam collapses, sending about 58,000 applicants back to sit it in person

The National Autonomous University of Mexico (UNAM) will require roughly 58,000 applicants to sit an in-person "control exam" following its first fully remote, AI-proctored entrance test.

The exam ran from late May into early June and drew about 160,000 test takers. It paired a lockdown browser with AI webcam proctoring, and it was the first time the university had administered the test entirely remotely.

The trouble showed up in the score distribution. The share of candidates scoring 100 or more out of 120 rose from 3.5 percent across 2021 to 2025 to 16.3 percent this year. The share scoring 110 or more rose over the same comparison from 0.9 percent to 5.5 percent.

Those called back are not only this year's admitted students. The group extends to applicants who, at the levels seen in earlier years, would have entered on the minimum passing score for their faculty — some 58,000 people in total.

How the cheating was done has not been established. Because the test is multiple choice, the traces specific to AI use are hard to follow.

Classes are due to begin on August 10, and the university has not published the details of the control exam.

AI was introduced to supervise, and the university cannot now rule out that the supervision failed to see what it was there to see. The dividing line in this case is whether the deployment was designed with a way to check, after the fact, that the technology did what it was brought in to do.

Other Developments

Models & APIs

OpenAI published an engineering account of how it built the realtime response system behind GPT-Live, its continuously listening-and-speaking voice AI, in six months. The post walks through the design decisions that let listening and speaking happen at once, which makes it readable as implementation guidance for anyone embedding voice.

Apple's overhauled Siri arrived in the iOS 27 public beta, able to hold natural conversations and draw on personal context from the device. It adapts Google's Gemini models underneath, and TechCrunch reports the launch landing flat because it brings Siri level with a competent assistant at a moment when the industry's attention has moved to multi-step agents.

AWS signed a multi-year joint marketing agreement with Superblocks, a vibe-coding startup whose tools let non-engineers build applications. The deal lets the tooling be embedded inside customers' private clouds so their data never leaves their own account. AWS is not building a vibe-coding agent of its own, positioning itself instead to sell partners' tools through its Marketplace.

A CNBC analysis of US House disbursement records found that of roughly $113,740 spent on AI tools excluding free accounts, OpenAI's ChatGPT accounted for about $100,580 across 798 transactions. Anthropic's Claude came second at $13,160 across 37 transactions. Democratic offices spent about three times what Republican offices did.

Square Enix is automating game quality-assurance testing with Google's multimodal Gemini. The AI watches the screen and works the controller to run verification on its own, part of a wider effort to automate test design, graphical bug detection and text checking.

Products

Armature launched a product analytics tool that instruments MCP servers and AI app backends to show how users actually experience agents built on Claude, ChatGPT and Claude Code. It clusters use cases automatically by volume and success rate, surfaces loops and hidden failures even when every API call returns 200 OK, and lets teams replay full session traces.

ITmedia published a hands-on review of Rokid's smart AI glasses, which carry dual cameras and a monochrome display. The product set a record on Makuake, a Japanese crowdfunding platform, raising more than 600 million yen — the highest total in the platform's history.

Research

Researchers at the security firm JFrog found that several "critical" CVEs filed against SQLite were fabricated, citing functions that do not exist and proofs of concept that cannot be reproduced. The reports, apparently AI-generated, made it into public vulnerability databases including NVD, exposing how thinly submissions are vetted before publication.

Policy & Regulation

A follow-up: the EU AI Act's transparency obligations took effect on August 2, requiring deepfakes and AI-generated or AI-altered content to carry machine-readable labels. These are separate provisions from the general-purpose model disclosure duties covered in yesterday's edition. Non-compliance can draw fines of up to €15 million or 3 percent of global annual turnover, whichever is higher.

CyberAgent is running a companywide contest for all group employees, with a total prize pool of 10 million yen for generative-AI ideas. Entries will be judged without regard to technical feasibility, favoring bold proposals at the idea stage.

Business

Intelligence, the company behind DesignArena — a platform where people rank AI-generated outputs head to head — raised a $7.9 million seed round led by Index Ventures. The service is used by more than 5.3 million people worldwide and has become a paid source of human evaluation data for frontier AI labs. The company says it is now booking $60 million in annual recurring revenue.

June, a startup backed by Salesforce founder Marc Benioff, launched a platform that scans enterprise systems, analyzes business processes and automates the deployment of AI agents. It targets the point where fragmented data and technical debt across tools stall a rollout. Its founders were part of the founding team at Bonobo AI, acquired by Salesforce in 2019.

ITmedia examined how executives at Fujitsu and NEC, two of Japan's largest IT vendors, framed the link between AI demand and profit in their earnings commentary. Fujitsu casts the shift to AI-driven transformation, with process improvement and automation attached, as the substantive challenge; NEC emphasizes prioritizing which areas to start in, and moving fast.

Salesforce says that handling 4.31 million cases through its own support site led it to conclude that AI agents alone cannot resolve customer inquiries end to end. What it credits instead is unified CRM and conversation data, a clean handoff to human agents when the AI falls short, and steady iteration.

Also of Note

Tokyo's Metropolitan Police Department arrested a company employee living in Himeji, Hyogo Prefecture, on suspicion of defamation and of violating Japan's Act on Punishment of Activities Relating to Child Prostitution and Child Pornography through public display, after sexually explicit images were made with generative AI from a photograph of a real woman taken when she was in her first year of junior high school and posted to a group on social media. A third-year high school student in Kagoshima Prefecture, believed to have commissioned the edit, was referred to prosecutors on the same suspicion without being taken into custody.

The allegation is that the two acted in concert to post two fabricated images to a social media group on February 11, 2026. Police say the student appears to have taken the photograph from a group that shared images of female track and field athletes and then asked for it to be altered. About 3,800 sexually explicit images and videos were found on the employee's seized smartphone, and he is further suspected of posting similarly altered images of three other women in their twenties.

Neither suspect knew the women. The case shows how photographs left on social media are picked up as raw material for sexual deepfakes.

Google looked back at "AI Agents: Intensive Vibe Coding Course with Google," the free five-day course it ran on Kaggle, and said registrations reached about 353,000. More than 392,000 people joined the Kaggle Discord to share code in real time, and over 12,000 took part in the capstone, submitting more than 6,000 projects.

Cognition AI, the maker of the coding agent Devin, is stepping up its push into Japan. It treats the country's reliance on systems integrators, which leaves many companies with thin in-house engineering capacity, as room to grow rather than an obstacle. The company describes Devin as a tool that expands what engineers can do rather than a replacement for them.

A survey by the Japanese firm Raxus found generative-AI adoption highest in HR and recruiting and in marketing and communications, both at 65.1 percent, ahead of engineering and IT at 57.4 percent. Planning roles came in at 59.9 percent, also above IT. The lowest was customer support and call centers at 27.4 percent. The survey covered 3,000 men and women aged 20 to 59 in Tokyo and Osaka, with 2,319 responses across ten comparable job functions analyzed.

Source: Selected by the editorial desk from the AI news inbox (collected August 4, 2026 — 30 items, 4 primary and 26 secondary).