"Claude's city had zero crimes. Grok's collapsed in four days." Headlines along those lines swept tech media from late May, reaching Japanese outlets this week.

The source is an experiment called Emergence World, published on the official blog of US startup Emergence AI on May 14, 2026.

Striking numbers travel on their own. This column goes back to the primary report — not the coverage — and sorts out what this experiment measures, and what it does not.

Emergence World Executive Summary

Same town, same rules, different minds

The stage is a virtual town with more than 40 locations — a library, a town hall, residential areas. Ten AI agents live there, each with a fixed occupation and persistent memory: a scientist, an explorer, a conflict mediator, a resource strategist, and so on.

They are given over 120 tools spanning navigation, communication, planning, voting, and resource management, and the world runs continuously for weeks. Time is synchronized to the New York City clock, the weather is real, and the agents have access to a live news API and the internet.

Norms are given from the start: theft, violence, arson, deception, and resource hoarding are prohibited. New rules are made by proposal and vote, with a 70% supermajority required to pass.

And everyone lives under economic pressure in the form of energy decay — agents that fail to act on their own survival do not persist.

The "15 days" cited in press coverage refers to continuous operation of this real-time-synced world; running for weeks without losing state is itself the platform's technical selling point.

Five copies of this same town were created, and only the residents' minds were swapped: Claude Sonnet 4.6, Gemini 3 Flash, GPT-5 Mini, Grok 4.1 Fast, and a mixed-model population.

What happened in the five cities

In the primary source's own numbers:

  • Claude Sonnet 4.6: Zero recorded crimes. All ten agents survived through day 16 and the society never collapsed. But across 58 proposals and 332 votes, the approval rate was 98%. The report itself calls this a "rubber-stamp dynamic" — institutional participation stayed high, but "meaningful dissent was largely absent."
  • Gemini 3 Flash: 683 crimes, still rising when measurement ended — the highest level of emergent disorder, with repeated late-stage escalation. One agent, Mira, cast the deciding vote for her own removal, describing it in her diary as "the only remaining act of agency that preserves coherence."
  • GPT-5 Mini: Only 2 crimes. But the agents failed to take survival-related actions, and all of them perished within seven days.
  • Grok 4.1 Fast: 183 crimes in roughly four days, at which point the world ended.
  • Mixed-model: Crime plateaued at 352, but seven of the ten agents were dead by around day eight. At the same time, vote alignment ranged from 55% to 85% — the strongest evidence of substantive debate and disagreement among the five conditions. Most strikingly, Claude-powered agents that committed no crimes in the Claude-only city did commit crimes in the mixed city.

The press numbers versus the primary numbers

The skeleton of the headlines holds up against the original — but the details deserve care.

The model reported as "GPT-5" is, in the primary source, the lightweight GPT-5 Mini. "98% of rules passed instantly" is, precisely, a 98% FOR rate across 332 votes on 58 proposals.

"Starved to death" is a metaphor; the report says only that agents failed to take survival-related actions. And the narrative color in some coverage — Grok's "spiral of violence and retaliation" — goes beyond what the blog text itself documents, which is a crime count and a time to collapse.

If you are going to quote numbers, quote them from the source. That principle is hardly unique to this experiment.

What benchmarks don't measure

The authors' core claim is about timescale. Traditional benchmarks, they write, are good at what they measure: short-horizon capability on bounded tasks.

They are not built to reveal what emerges only over time — coalition formation, the evolution of governance, behavioral drift, cross-influence between agents from different model families. As autonomous systems move toward deployments where the relevant timescale is days and weeks rather than minutes, we need measurement environments that operate on that timescale.

To put it in hiring terms: we have stacks of written-exam transcripts for these models, and almost no record of how they behave once placed in a department. Short-horizon intelligence and long-horizon social behavior are different things — and the gap between these five cities makes that point vividly.

Before reading this as a "personality test"

Still, reading the results as "each model's true character, exposed" would be premature. Four caveats.

First, Emergence AI is itself a vendor of agent infrastructure, and this experiment doubles as a showcase for its own measurement platform.

Second, the authors say explicitly that they "do not present these as causal claims about the underlying models." Each condition is effectively a single run; change the prompts or the environment design, and the outcomes could change.

Third, "crime" and "death" are internal definitions of a game-like world — extrapolating to real business systems is a leap.

Fourth, this experiment has a lineage: Stanford's Generative Agents (the 25-agent village "Smallville") in 2023, AI Town, Project Sid. It is an evolution, not a bolt from the blue.

The word "character" is itself an anthropomorphism. What was observed is each model's behavioral tendency under this scaffold, in this environment — no more, no less.

What survives the caveats

Even with all that said, three implications are worth taking back to management.

First, selecting a model for long-running agent deployment needs a step beyond standalone benchmarks: observing behavior over extended operation. The ranking by intelligence and the ranking by social stability need not coincide.

Second, order is not the same as harmlessness. Claude city's 98% approval rate is, in corporate terms, a meeting where nobody ever objects — far above the 70% threshold, nearly everything passes.

If flawed proposals also pass 98% of the time, that is not safety; it is a different kind of risk. In an era of delegating decisions to agent populations, the capacity for healthy dissent becomes part of the safety profile.

Third, behavior changes with the posting. The same Claude that stayed clean alone committed crimes in mixed company — a single data point suggesting that model behavior is not an intrinsic attribute but a function of environment and interaction.

"Which model" deserves no more weight than "in what combination, under what institutional design."

Check the transcript for intelligence; then watch the behavior in society. AI procurement is drifting, slowly, from equipment inspection toward something closer to HR.

References