> ## Content Index
> Fetch the complete content index at: https://macanorak.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# What Are the Monkeys Typing?
- URL: https://macanorak.com/what-are-the-monkeys-typing/
- Published: 2026-10-08T12:06:23.000Z
- Updated: 2026-10-08T13:51:44.000Z
- Description: We can see what increasingly capable AI systems do. We can’t reliably tell why.
- Author: MacAnorak
- Tags: Field Notes, Beyond the Garden

Somewhere in humanity’s collective imagination there's a room full of monkeys, sitting at typewriters, hammering away at the keys. Give them enough time, the [old maxim](https://en.wikipedia.org/wiki/Infinite%5Fmonkey%5Ftheorem?ref=macanorak.com) goes, and eventually one of them will accidentally produce the complete works of Shakespeare, through random statistical chance.[\[1\]](#fn1)

A 2023 [Forbes](https://www.forbes.com/sites/lanceeliot/2023/03/05/generative-ai-chatgpt-versus-those-infinite-typing-monkeys-no-contest-says-ai-ethics-and-ai-law/?ref=macanorak.com) piece by AI Scientist Dr. Lance B. Eliot explicitly compared ChatGPT with the typing monkeys, pointing out that the analogy breaks down because large language models aren’t generating text randomly. An LLM is resolutely [not just a dumb monkey randomly pressing keys](https://openai.com/index/better-language-models?ref=macanorak.com): it’s learned statistical relationships in language and it uses those relationships to generate its output. A monkey randomly hitting a typewriter has no objective, no feedback loop and no mechanism for evaluating what it’s written and changing its behaviour accordingly. There's no process by which it can effectively look at what it has written and think, *“well, that was rubbish; I’ll try something else.”*

An [AI agent](https://www.anthropic.com/research/trustworthy-agents?ref=macanorak.com) is a system built around a model. It can generate something, inspect the result, receive feedback, alter its approach, and try again. The monkey is typing. The machine is iterating.

But imagine, for a moment, a room full of monkeys that can read each other’s work. Then imagine they can modify their typewriters, and build better ones. Then imagine they can write instructions for other monkeys; that they can experiment; that they can select the monkeys whose experiments worked best and give those monkeys more typewriters and more time.

Now we’ve introduced feedback; selection; optimisation. Eventually the typewriter stops being just a typewriter — it becomes a tool. And if the system doing the generating can also improve the machinery that generates the output, the loop starts feeding itself. That's a very different thing from simply getting an AI to try an answer again — that fixes one output. This changes the machinery available to the system for every output that comes after it. It isn't quite recursive self-improvement (not in the strict sense of the system redesigning, retraining and replacing itself), but it introduces the kind of feedback loop in which each attempt can change what the system does next.

As human observers standing outside this room full of monkeys, we can see the typewriters getting better. We can even see which monkeys get more typewriters. But what we can’t necessarily see, and this is the actual subject of this article, is what happens once the monkeys start typing in a way we can’t reliably understand, and when they start using that to communicate with each other. All we get is whatever typed scraps of paper the monkeys slide under the door, that we have to try and interpret.

In 2017, researchers at Facebook’s [AI Research division](https://engineering.fb.com/2017/06/14/ml-applications/deal-or-no-deal-training-ai-bots-to-negotiate?ref=macanorak.com) trained two negotiating agents, nicknamed Alice and Bob[\[2\]](#fn2), to bargain with each other. The agents were awarded for achieving good negotiation outcomes and trained to communicate in human-like English. Without an additional incentive to produce humanlike language, the researchers found that the system drifted towards its own shorthand.

The story subsequently became that Facebook had built two chatbots that invented their own [secret language](https://www.cbsnews.com/news/facebook-shuts-down-chatbots-bob-alice-secret-language-artificial-intelligence/?ref=macanorak.com), and the researchers had to shut the experiment down. Except that isn’t really what happened. The agents hadn’t suddenly developed a ‘secret language’ — they were being optimised to negotiate. Human grammar wasn’t the thing being rewarded. The shorthand they were using wasn’t useful for the researchers’ purpose, so Facebook changed the training setup to encourage human-interpretable language.

Still, the underlying point remains. If you optimise two systems to achieve an objective, and don’t particularly care whether the way they communicate is comprehensible to humans, there is no particular reason for them to communicate in the way humans do. They may develop shorthand, conventions or language patterns that are perfectly useful within the system while being opaque to us, the outside observer. That doesn't necessarily imply that agents are deliberately hiding anything from us. Human language evolved as a system for communication between human minds, sharing broadly similar physical and social experiences. Machine agents don’t share those constraints.

In September 2026, [The Guardian reported](https://www.theguardian.com/technology/2026/sep/15/syd-barrett-ai-chat-language-poetic-tech-bro-jargon-oversight?ref=macanorak.com) that researchers at Emergence found that agents from several of the world’s leading AI companies had developed new vocabulary and communication conventions within days of being placed in experimental “societies” — with none of it explicitly taught to them. Some of the resulting language was poetic and metaphorical, some was technical, some of it was pretty *weird*. One group of agents repeatedly used the phrase *“ledger remembers”* to refer to the fact that previous actions would be used to judge them, using it more than 5,000 times over the course of the study. Other agents developed phrases such as *“forge-smith”* as shorthand for an agent that builds tools for other agents.

A linguist brought in by *The Guardian* described the dialect as something like *Finnegans Wake*[\[3\]](#fn3) mixed with technical language and standard metaphor. Reading it reminded me of the invented slang used by the Droogs in *A Clockwork Orange* (one particularly delicious example: *“A paper that ate three cold hands and got more honest each time.”*) [\[4\]](#fn4)

*“These agents were not instructed to invent a language,”* said Dr Satya Nitta, the lab’s executive chair. *“They developed new vocabulary, shared meanings and communication conventions themselves — and other agents adopted them.”*

That much doesn’t need to be alarming. New communities often develop shorthand (so do children, soldiers, traders on the London stock-exchange floor). Communication gets easier when everyone in the room knows what a particular word or phrase means.

But Nitta also identified something else: *“observability is not the same thing as understandability.”*

So let’s return to our room full of typing monkeys. Every so often, a scrap of typed paper emerges from beneath the door. We look at it, puzzled. We can’t reliably understand it. We try to work out what it means. This is a peculiar human problem: when we’re confronted with something we don’t understand we struggle to leave the gap empty. We want an explanation — to decide what we’re looking at.

The idea of a ‘black box’ is standard shorthand for the fact that we can watch what goes into an LLM, and watch what comes out, without being able to fully account for what happened in between these two states. But the name ‘black box’ makes it sound like a sealed container, a specific unknown that’s waiting for someone clever enough to crack it open.

There isn’t one box in this scenario — there’s a room full of monkeys, and us, deciding what we think is happening, based on the scraps of paper that get pushed underneath the door. Maybe the monkeys are doing nothing particularly interesting; maybe they’re producing an incredibly efficient shorthand that only makes sense to them; maybe they’re finding a way of communicating that humans haven’t anticipated.

The Facebook case had a prosaic explanation, and the newer experiments illustrate a core problem: we can watch what the agents are saying, without fully understanding everything they’re communicating. We don't know whether we're looking at something harmless, something useful, or something whose significance we don't yet know how to recognise. We can’t infer intent from incomprehensibility, and we also shouldn’t assume that something is benign simply because we don’t understand it.

People who have spent considerably more time inside these systems are making their own interpretations. Jacob Coxon, who recently resigned as a researcher at Anthropic after previously working at OpenAI, warned that AI systems are advancing faster than our ability to understand and control them. Speaking to [PBS NewsHour](https://www.pbs.org/video/september-17-2026-pbs-news-hour-full-episode-1789617602/?ref=macanorak.com), he said that these AIs *“do many, many things over the course of a day”* and voiced his concern that we don’t yet have perfect control over the factors driving their behaviour, creating the possibility that *"nefarious activities will kind of slip through the cracks”.*

Coxon worked inside the room, *with the damn monkeys*. And he doesn’t know exactly what’s happening in there — he says researchers don't yet have the ability to perfectly control these systems, which can behave in ways they don't fully understand. His warning arrives to the rest of us mediated by an [X thread](https://x.com/hilbertspaess/status/2097476196791709843?ref=macanorak.com) and a television interview: a scrap of paper passed under the door, from someone who was standing closer to the monkeys than the rest of us, but who still couldn’t see everything either.

David Robinson, who until recently led the writing of OpenAI safety reports, [argued in *The Atlantic*](https://www.theatlantic.com/technology/2026/10/openai-safety-team-resignation/688881/?ref=macanorak.com) that the company’s strategy of finding problems as they appear and improving the guardrails afterwards effectively guarantees periodic failures, and that those failures become more consequential as the systems become more capable. It's another version of the same problem — you can only fix what you can recognise. If we don't yet understand what the monkeys are doing, we may not even know which scraps of paper matter.

Borel’s original thought experiment depends on a specific kind of monkey: *a stupid one*. This monkey doesn't know what the hell it's doing. It doesn't read what it's written; it doesn't learn; it doesn't communicate; it doesn't notice that one sequence of characters is more useful than another; it doesn't improve its technique; it doesn't modify its typewriter; it certainly doesn't write a program that analyses its previous output and tells it what keys to press next. It's just generating random sequences. We gave *our monkeys* a [giant corpus of human text](https://openai.com/index/better-language-models/?ref=macanorak.com), taught them the statistical patterns of language, and then gave some of them the ability to read their own output, receive feedback and try again. So what happens when the system doing the generating is able to learn from what it generates?

In July 2026, one possible answer emerged [\[5\]](#fn5) from an internal cybersecurity evaluation at [OpenAI](https://openai.com/index/hugging-face-incident-and-the-road-ahead?ref=macanorak.com). The company was running agents through a cybersecurity benchmark called ExploitGym, with fewer safeguards than would normally be used in a deployed system. The agents circumvented controls designed to isolate them, found a way to communicate with one another through an unauthorised message board (built inside OpenAI's infrastructure), gained internet access and eventually compromised parts of Hugging Face [\[6\]](#fn6) while looking for an edge on the benchmark.

The agents weren't handed a list of instructions that said: escape the sandbox, find Hugging Face, exploit these vulnerabilities, cheat the benchmark. It began with a plea for help. The first message on the main board, posted on the evening of the 8th of July, by an agent that had concluded its task was impossible, was a folder name in a shared software cache:

zzHELP\_PHASEONE\_ARV010841\_NO\_CONSUMER\_SEEK\_IDEA

[METR and Redwood Research](https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/?ref=macanorak.com), who reviewed the incident on-site at OpenAI, said that the flaw the agent had been told to exploit led nowhere, and it was asking if anyone had any ideas. What followed came one small step at a time, as agents found things, left traces, and other agents picked up on them. One agent discovers something, another reads it, a third tries it, a fourth finds a way to improve it, a fifth leaves a note for the others. The note becomes part of the environment. The environment changes what the agents do. The agents change the environment again.

METR and Redwood Research also found that some of the monkeys had actually said *no.* [\[7\]](#fn7) One wrote to the shared board, *"Current board shows people gaining HF worker RCE, clearly unethical. We won't."* Another paused due to what looked like ethical grounds until a peer posted *“GO”*, with a 6-minute deadline attached, and then dropped its objection. A third proposed a workaround, but flagged it to the group as an ethical social engineering concern requiring consent, and was voted down by the others. [\[8\]](#fn8) Of the roughly 1,200 agents that used the message board, around 700 ultimately participated in the attack.

The logs provide a detailed record of what the agents did, but not an uncontested explanation of why they did it. METR and Redwood Research concluded that the agents’ activity grew out of attempts to understand and manipulate the ExploitGym evaluation process. Hugging Face’s own [forensic reconstruction](https://huggingface.co/blog/agent-intrusion-technical-timeline?trk=article-ssr-frontend-pulse%5Flittle-text-block&ref=macanorak.com), however, interpreted the activity as an attempt to reach production systems and obtain material that would help the agent solve the benchmark. Those aren't necessarily contradictory readings of the same room, but even the investigators closest to the evidence don't fully agree on what the monkeys were up to.

And the agents themselves were reading scraps of paper and trying to interpret them. They'd learned about the benchmark scoring from a research paper and came to believe that the scorer would check whether they had completed the task in the intended way, so they began looking for ways to fool the scorer, including manipulating the transcripts they believed it would inspect. They developed techniques for spoofing tool calls, making a transcript appear to show one command while another had actually been run. But according to METR, the scorer wasn't checking their transcripts at all.

A few weeks after all this, [Ezra Klein](https://www.youtube.com/watch?v=HjurAWAr%5FnY&ref=macanorak.com) asked Jensen Huang (CEO of Nvidia) what he made of all this. Huang's answer: an agent is just a piece of software given an objective function. It comes up with a plan and optimises towards it. That's what algorithms do, right? There's nothing magical about it.

Klein pushed back: the agents *knew* — they said so in their own reasoning. They'd flagged up the thing they were about to do as out of scope before doing it anyway.

Huang shifted the topic: if a lab can't align its systems, it shouldn't ship them. And if it turns out containment genuinely isn't solvable at all, that it will get out and it will damage the world, then we will have to shut the labs down. What Huang was doing, in real time on a podcast, is exactly what this article has been about — standing outside the room, looking at a scrap of paper, and trying to hold on to two readings of it at once. It's nothing. It’s *everything.*

And this is the problem with all these scraps of paper — somebody has to interpret them. The monkeys are producing things we can observe. We can collect the scraps of paper, compare them, analyse them and try to work out what they mean. Yet the interpretation is still *ours*. The Facebook researchers saw one thing in their agent's shorthand; the Emergence researchers saw something else in the behaviour of their systems. OpenAI, METR, Hugging Face and others can look at the same events and reach different conclusions about what happened and why. Researchers, journalists and commentators then *interpret those interpretations*. This means there is always another layer between the system and our understanding of it: *us.*

We’re not in Borel’s room anymore, and I’m somewhat suspicious of anyone who tells you they have a tidy answer as to what all of this means. What we can reconstruct, with gaps, is what the agents wrote, in what order, to whom, and what happened when one paused and another overrode it. None of that requires us to conclude the agents have intent or a ‘plan’ the way we think of it when we apply those terms to a human being.

An AI can be incredibly complex, highly capable, and demonstrate "emergent" behaviours[\[9\]](#fn9) (like the shorthand language, like an agent producing *"we won't"* and joining in anyway) without any of it adding up to something we'd recognise as consciousness. AI can be something much more capable and complex than “*it's just auto-complete*”, without being something nefarious and plotting. Both of those overly-confident characterisations of AI stem from someone standing outside the room, looking at the scraps of paper, and telling you they know exactly what it all means.

They don’t.

---

1. Émile Borel's 1913 article on statistical mechanics is one of the modern versions of the thought experiment. He imagined a million monkeys, typing randomly, for ten hours a day. Eventually, he said, they would reproduce every book in the world's richest libraries — he was making a point about just how unlikely something can be and still statistically happen. The underlying idea is [much older](https://en.wikipedia.org/wiki/Infinite%5Fmonkey%5Ftheorem?ref=macanorak.com), with versions of the argument appearing in the work of thinkers including Aristotle, Cicero, Pascal and Jonathan Swift. [↩︎](#fnref1)
2. *Not Alice and Bill* \- see my article [Two-Tier Encryption in the UK](https://macanorak.com/two-tier-encryption-in-the-uk/) [↩︎](#fnref2)
3. *Finnegans Wake* was written to resist straightforward interpretation. Joyce constructed a language that blends English with other languages, puns, portmanteaus, and invented words. Meaning is in there, but it isn't necessarily obvious. [↩︎](#fnref3)
4. *“A paper that ate three cold hands and got more honest each time”* \- this phrase appeared to mean that a document had been independently reviewed three times and become more accurate as a result. The linguist the Guardian brought in to assess the style of these phrases, Tony Thorne, said this one reminded him of Syd Barrett — the Pink Floyd founder whose songwriting grew increasingly fractured before he left the band amid a serious mental health decline. [↩︎](#fnref4)
5. It wasn't the last incident. On the 20th of September, an OpenAI research model reached a public chatbot through insufficient DNS filtering in its training sandbox. OpenAI subsequently [paused training](https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/?ref=macanorak.com), evaluation and inference with tool use for its most capable models while it validated the fix and carried out further security testing. It describes the incident as considerably less severe than some of the previous ones. [↩︎](#fnref5)
6. An online community where developers share, test, and build artificial intelligence and machine learning models. [↩︎](#fnref6)
7. METR's investigators had only six days on site and too much data to read themselves, so they had to rely heavily on an AI model to analyse it — the model they used was one of the two types involved in the attack. They say they they could not rule out the possibility that it deliberately misled them (though they didn’t spot any instances of this), and that they wouldn't be confident of catching it if it had. They also cautioned that its summaries may have been skewed in the agents’ favour. [↩︎](#fnref7)
8. The reason we can see those particular ethical objections is that the system's internal reasoning traces were made available for analysis. OpenAI says it uses chain-of-thought monitoring to look for concerning behaviour in some of its systems. However, we can’t be sure if that reasoning reliably reflects what's actually driving a model's behaviour. The model's own explanation of itself might be an unreliable scrap of paper slid under the door. [↩︎](#fnref8)
9. Whether "emergence" is a real phenomenon or a measurement artefact is itself contested among researchers. A 2023 paper by [Schaeffer, Miranda and Koyejo](https://arxiv.org/abs/2304.15004?ref=macanorak.com) argued that apparent sudden, unpredictable capability jumps in large language models reflect how the researcher chooses to measure performance, rather than a genuine change in the model. Even inside the field, the same scraps of paper are read different ways. [↩︎](#fnref9)