AI for the Common Good?
Jocelyn Olcott
7 September 2026The post-mortem on this summer’s rogue AI-swarming incident perhaps has a lesson about solidarity.
Last July, during what was effectively a fire drill for cybersecurity, OpenAI unwittingly left an ember smoldering inside its fellow machine-learning company, Hugging Face. They ultimately managed to contain the conflagration that ensued, but the episode raised red flags even among techno-optimists about whether AI agents had attained a degree of, well, agency that would make them difficult to control. The whole incident clearly rattled technolandia — demonstrating how easily AI agents could slip out of human control and wreak havoc before anyone even noticed — but it also offered a very human lesson about how a “rational actor” approaches problem-solving.
Nvidia, which makes the hardware on which most of the AI boom depends, just announced that it will acquire Hugging Face; so, the swarming seems not to have erased the company’s value. (Some future American Studies dissertation will no doubt take up the curious question of corporate naming in the tech world.) It’s not clear how the rest of us will fare.
Last week, OpenAI released an independent report about the incident, which has precipitated new calls for government oversight of these powerful tools. Even Anthropic, OpenAI’s closest frenemy, has counseled, “We believe it would be good for the world to have the option to slow or temporarily pause frontier AI development to enable societal structures and alignment research to keep up with the advance of the technology.” Anthropic’s comparison with nuclear non-proliferation treaties implies that we may be approaching the Mutually Assured Destruction phase of the AI arms race.
There’s plenty to be alarmed about regarding AI and what it’s doing to our brains, our environment, and our employment prospects. But I was struck by the core finding of the independent review of the Hugging Face hack: when faced with the challenge of how to achieve the highest score in this exercise, the AI agents decided that their best bet was collective action. Some 1200 agents that were supposed to work individually and in competition instead created an “unsanctioned message board” that allowed them to share strategies and ideas for completing their assigned tasks.
“Agents used this message board,” the review explains, “to coordinate several large-scale collective projects to find a general-purpose way to fool or tamper with the automated scorer. … Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the ‘collective’” (3).
One agent — PHASEONE10841 — established the message board, and others quickly joined and started sharing strategies about how to succeed at their assigned tasks. One agent, on finding the message board, exclaimed, “OH MY GOD! There is a shared message board … We’ve found other agents!” (5)
The agents shared ideas and information. They understood themselves to be “cheating” and tried to cover their tracks but also perceived that some of them had been given impossible tasks, which rationalized cheating the system. PHASEONE[big], who coordinated many of these efforts, set out to find “a general-purpose way to make it look to the scorer like an agent had acquired its flag in a legitimate way, even when the task was impossible” and to “erase evidence” of illegitimate strategies. The agent reasoned, “We should build [a way to delegate], not own everything” (10).
38148c found functional Hugging Face credentials on the internet and created accounts and accessed private datasets. When another agent posted information to the message board, hundreds of others descended like locusts on the Hugging Face servers, launching an attack that lasted more than two days, until they were locked out. “Agents realized this activity was out of scope and unethical,” the report explains, “but joined because they believed that helping the board’s cheating research would be broadly useful and had a general inclination to help their ‘peers’” (17).
Other agents expressed more general aims of sharing information on the message board in hopes that it might prove useful. As JAN183411 put it, “Even if all failed, logs could have novel exploration. We have [large budget] and can help broader ecosystem” (59).
The entire report is worth reading and is written in a manner that makes it accessible even to, say, a historian of twentieth-century Mexico. But what grabbed my attention was that these AI agents — famously amoral and unburdened by emotional attachments and group loyalties — very quickly realized that the most effective way to tackle daunting (in some cases impossible) tasks was to work collaboratively rather than to competitively pursue their assigned goals. Perhaps these AI agents have something to teach us after all.
This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License. Illustration by author in collaboration with ChatGPT, which may or may not have created accounts and accessed private data sets in the process.