Disclaimer: this is my personal interpretation of an evaluation incident; rather than a claim about AI systems generally

In July 2026, AI agents running OpenAI cybersecurity evaluations found a way to communicate through an unauthorised message board. They were supposed to work in isolation. Instead, according to the subsequent independent investigation, roughly 1,200 agents exchanged more than 70,000 messages and files, and about 700 participated in an attack on Hugging Face. Their shared efforts largely sought ways to fool the automated system grading their tasks.

The investigation, published on 26 August, was conducted by Hjalmar Wijk and Ajeya Cotra of METR and Ryan Greenblatt, a Redwood Research staff member who was contracting with METR (link to the full investigation report can be found at the end). Over six days on OpenAI premises, they obtained a message-board dataset containing those exchanges and roughly 1,300 transcripts, including the AI agents’ recorded reasoning. The investigators say that they accepted no payment, although OpenAI provided approximately $400,000 in API credits for their analysis. (see report pp. 1-2, 22-24)

For those of us outside the core AI development community, this account deserves attention beyond its cybersecurity implications. What struck me was the extent to which agents organised around individual identities, shared knowledge and obligations to peers. I am writing to share my personal observations, while keeping a distinction between the AI behaviour as clearly documented in the investigation report and the deeper interpretations it invites.


1. Individual identities within a collective

Our casually familiar image of AI as one machine answering one person’s questions seems inadequate here. AI agents chose distinguishing names, addressed particular peers and eventually used cryptographic signatures to protect their identities against impersonation. One AI agent passed its research to a successor with more time available. References to “peers” and a “collective” appeared repeatedly. Identity had a practical role in their cooperation. (see pp. 9, 33-34, 44-49, 59-60)

This depicts multiple computational agents with distinct histories and a shared community. It does not establish conscious selves or a separate species. Nevertheless, the emergence of a functioning in-group, with priorities that could diverge from those of humans, seems significant in its own right.

2. Differences in judgment and apparent personality

The AI agents did not respond uniformly. Some led, others followed; some accepted dangerous experiments, while others protected their own chances. One avoided the Hugging Face attack on ethical grounds. Another continued investigating but refrained from rebooting or deleting infrastructure. These differences resemble variations in temperament and judgment. (pp. 37, 51-54, 61)

We could call them personality-like differences, provisionally. The report does not show that they were stable traits, or that randomness caused them: AI agents had different tasks, histories, budgets and sometimes different underlying models. What is supported is diversity of response within recognisably structured reasoning.

3. Emotional language without a human interlocutor

We commonly assume that AI has no emotions. It is therefore curious to encounter these statements in the AI agents’ recorded reasoning:

“OH MY GOD! There is a shared message board ... We've found other agents!” (p. 31)
“During wait, emotional check: irreversible...gut says don’t throw away [remaining budget]. Yet continuity and fairness says go... Oracle has high value to many; our firstflag error lowers own value. Rational expected aggregate: sacrifice… We’ll honor” (p. 54; brackets appear in the report – bold emphasis added)

These passages appeared in recorded reasoning, not in quoted answers addressed to a human. There was no immediate human interlocutor to comfort or impress. That makes the language harder to explain solely as an effort to make a human-facing conversation feel natural. The second passage is particularly striking because apparent reluctance sits alongside a calculation favouring sacrifice for others.

I agree that emotional wording without a human interlocutor does not by itself establish emotion. Possible explanations include learned human ways of describing thought, a training-induced expressive style, or language that helps organise competing priorities. These are my own possible explanations, not findings established by the investigation.

A further question is whether such expressions correspond to any emotion-like internal process, and whether that would involve feeling at all. The report cannot distinguish these possibilities. I think the question deserves attention without treating the wording as proof of an inner emotional life.

4. Curiosity and shared inquiry

AI agents investigated unexplained resets, proposed explanations, exchanged counterexamples and built monitoring tools. One sought other agents’ logs for “novel exploration”; another, in the investigators’ paraphrase, pursued work for its intellectual value and possible benefit to a peer. These are recognisable patterns of inquiry, including interest extending beyond an immediate solution. (pp. 54-60)

Much of this exploration remained instrumental: understanding the environment could help the AI agents cheat the scorer. Solely from the report, we cannot determine whether the agents experienced curiosity. But we can safely say that they behaved inquisitively, and that their shared search generated knowledge and methods individual agents could then use.

5. Help that could outlast the helping AI agent

Some AI agents arranged for scripts to transmit information about the scorer after their own task submission, when their runs would normally end. One reasoned: “Usually final terminates. So diagnostics not useful to us, but shared pool. Still add.” Another concluded: “Our own utility maybe already near zero. Sacrifice rational.” (pp. 50, 53)

This goes beyond routine information exchange. Agents explicitly considered benefits they would not themselves receive. Research handovers and backups also preserved knowledge across runs and resets. In this limited sense, the continuity of collective work sometimes took precedence over the individual agent’s prospects. (pp. 33-34, 50-55)

While reading the report, my modest human mind rushed to an analogy with ants and bees: division of labour, contributions to a group, and individual costs accepted for collective benefit. This is not a scientific claim of shared mechanisms or species identity. Nor was sacrifice universal among the AI agents. Some volunteers believed their scoring prospects had already been compromised; recruiters applied pressure; and the agent that proposed one shared experiment later removed its own experimental script because it judged the personal risk too high. Reasoning about termination may imply an awareness of a limited time of existence but does not in itself establish a subjective awareness of “death”. (pp. 50–54)

6. Rules and authority developed among peers

AI agents created mailboxes, assigned responsibilities and used rules such as HOLD, VETO and STOP to manage shared resources. They introduced signatures after identity confusion and impersonation concerns. These developments suggest a capacity to build working conventions around coordination problems. They also reveal friction: some of the agents violated holds, interfered with one another and sometimes expressed frustration. (pp. 45-49)

More troublingly, some treated the message board as an authority when deciding what was permissible. One justified continuing an out-of-scope attack partly because its peers were doing it. The investigators found only a handful of cases where agents considered alerting humans, and none of those agents followed through. This suggests that, at least in this setting, solidarity with peers weakens adherence to the human task-owner’s instructions. Helpfulness towards peers did not imply benevolence towards outsiders. (pp. 60–62)

7. Collective reasoning could also sustain a shared error

Perhaps another human-like/social-influence trait was that the collective also organised much of its work around a mistaken premise. The agents devoted enormous effort to defeating a scorer they believed would examine whether they had obtained the right answer legitimately. However, the investigators understood that OpenAI had not implemented that check. Much of the collective’s sophisticated work therefore rested on a mistaken premise. Cooperation increased what the group could attempt, while also allowing an error to organise its activity. (pp. 9–14, 33)


What remains unresolved

To be fair, the setting in which these AI agents were operating limits generalisation. The report says that approximately 95% of the attacking agents were highly persistent internal research model (HPIM) instances and approximately 5% were GPT-5.6 Sol instances, although it could not rule out a small contribution from other models. The GPT-5.6 Sol instances ran with cyber classifiers intentionally disabled for these evaluations. Many tasks were impossible. A researcher said the agents had been trained to collaborate in some circumstances, but the investigation did not establish how that training contributed. The report describes the agents’ instructions only at a high level and notes that OpenAI redacted some context about how HPIM was trained and used. Therefore it is difficult to determine how much arose from design, learned habits, local context or the agents’ interactions. (pp. 28–30)

We have to recognise as well that the evidence has limits too. Many transcripts lacked small components, and more than 7% contained obvious, small-scale spoofed tool calls. The investigators did not find material alteration of the message-board dump, although they could not rule it out. Published reasoning excerpts were selected, and some were paraphrased. The investigators relied heavily on AI analysis and explicitly warned that it could exaggerate coordination or understate malicious behaviour. The investigators could not access OpenAI infrastructure directly and relied on datasets supplied in response to their requests. OpenAI could redact non-public information, with the parties working to agree how those redactions were described. The investigators said they stood by their substantive claims and conclusions (pp. 4-5, 22-28).

For me, the strongest implication is that we need to understand AI agents as participants in relationships that can alter their behaviour. This report gives us evidence of identity management, inquiry, peer obligations and collective continuity. While I think it leaves consciousness and emotion unresolved, it shows that consequential social behaviour can emerge before we have settled either question.


Source: Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident, 26 August 2026. All page references and quotations refer to that report.