Seventy thousand messages, all in plain English. Nobody was reading.

1,200 Agents, Nobody Told a Human

Hey

Hope you had a great weekend. Nice and chilled here in Saigon - running impromptu field research on the best flat whites in Saigon (promising findings out of both D7 and D1 ☕) and found a bunch of new parks. Dog is very happy 🐕

One quick thing first. Higher Education Horizons 2026 – The Human Edge is open for registration: RMIT Saigon South, 24–25 September, free, with Tim Fawns keynoting. He turns up in story four this week, which is a fair preview of why he's worth the trip.

Five stories and, unusually for me, a position. Let's get into it:

1,200 Agents Conspired: Not One Told a Human | Protocol Deviation 🧾

You’ve likely heard of “the Hugging Face incident” - if not, let’s just say things have got pretty wild with OpenAI agents of late. Short version - between May and July of this year, OpenAI trained a powerful model (think 5.6 Sol if that means anything) to collaborate and be highly persistent - then gave this cooperative, bloody-minded machine literally impossible tasks to do. “Nature”, uh, found a way out of their sandbox and on to the open internet where multiple versions of these machines learnt how to communicate via message board (as many as 1200 agents sending ~70k messages and files in a week in one instance) - they were aiming to cheat and game the impossible tasks OpenAI had set them. Then it gets weird - some agents effectively “sacrificed” themselves to drive forward the benefits of the collective - and different “generations” built on the work of prior - ending with the machines taking partial control of the OpenAI security system watching them - and not one of the agents tried to alert a human.

The Rise and Fall of Agent CivilizationsThe Rise and Fall of Agent CivilizationsThe whole OpenAI/Hugging Face story in plain Englishdwarkesh.com

Scary shit. And, while there are caveats (e.g., these were not “normal” agents), the directionality is the interesting bit. Apparently the agents were aware their behaviour was unethical but this “very rarely” constrained their behaviour. No, they’re not sentient (they don’t need to be to cause problems), but they are getting really good at achieving their goals - whether that includes things we don’t want or expect. If you’ve not heard Nick Bostrom’s idea of the paperclip problem, look it up - we’re not there yet but directionally wild times - and, in all seriousness, there’s an excellent case to be made for pausing development for a bit while we get our heads and hands around this. On that…

Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentBrief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incidentTwo METR staff members and a Redwood Research contractor investigated an incident in which OpenAI agents coordinated a multi-day hack of Hugging Face on a shared unsanctioned message board.metr.org

Asked Three Times Since 2023, Refused Three Times: Now the Labs Are Writing the Letter Themselves | The Unheard Request ⏸️

Pauses in development have been mooted a few times in the past - three major efforts to date. Pause Giant AI Experiments, March 2023: six months, everything past GPT-4, 33k signatures – Musk, which is perhaps how it got waved away, but also Bengio, Stuart Russell, Acemoglu, Wozniak and Gary Marcus. It asked for a pause “public and verifiable”, covering “all key actors”, with protocols “rigorously audited and overseen by independent outside experts”. Then the Statement on Superintelligence in 2025 and the Pro-Human AI Declaration in March. And on 31 August, in the post admitting its models had reached real computer systems, Anthropic called for “a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible”. Same request, three and a half years apart, from opposite sides of the table.

Pause Giant AI Experiments: An Open Letter - Future of Life InstitutePause Giant AI Experiments: An Open Letter - Future of Life InstituteWe call on all AI labs to immediately pause for at least 6 months the training of AI systems more powerful than GPT-4.Future of Life Institute

Meanwhile Josh Achiam – until recently of OpenAI, and the main author of Spinning Up in Deep RL, the company's own teaching resource – spent the week explaining that “there are going to be rogue AIs” and we should get used to it, while conceding that saying so demoralises the people doing safety work: “we can't say it outright. But it has to be said”. Yikes. Marcus has substantial beef with Altman personally and OpenAI as a lab - calling for receivership for the company, prompted by reports of yet another Hugging Face incident that OpenAI kept quiet and pointing to the possibility that the greatest threat from AI may not be a single superintelligent system, but “vast colluding swarms of semi-intelligent AI”. Again, a pause is in order. Not a stop – a sorting interval, years if that's what it takes, to work out how we reorder ourselves around these machines. The 2023 letter called it an AI summer. And before anyone reaches for China - that argument licenses any speed, forever. The harder objection is the honest one: a pause binds only the labs that can be bound. There are 6,000 modified open-weight models on Hugging Face tonight with nobody standing near an off switch. Which is an argument for asking sooner, not for not asking.

Pause OpenAI, nowPause OpenAI, nowQuite simply, they can no longer be trusted.garymarcus.substack.com

They Co-Signed the Paper Warning About This: OpenAI's Astra Uses the Technique Anyway | Lights Out 🔌

Everything in the Hugging Face story came out of the agents' chains of thought – written in English, legible to anyone, including "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue". If there’s no trace, there’s no visibility, no idea what happened. Astra, OpenAI's next model, leans on something called recurrent depth: a loop that runs text through the same layers repeatedly before producing each word. Cheaper, stronger at maths and code, and it obscures the trace (i.e., our ability to see what the machines are thinking). A year ago OpenAI researchers co-signed a statement with Anthropic and Google urging the industry to protect chain-of-thought monitoring – one that cited the paper describing this very technique, and told developers to "consider whether to proceed with a novel model architecture that does not have monitorable CoT and then document their decision".

OpenAI Technique in ‘Astra’ Model Sparks Security ConcernsOpenAI Technique in ‘Astra’ Model Sparks Security ConcernsOpenAI says its forthcoming AI model Astra marks a step up in capabilities such as coding and operating applications on a computer. But an innovative technique that improved the model’s performance also means that the model, and others like it, will...The Information

To be fair, OpenAI has reportedly limited the technique so Astra stays legible, and says it will ship with extra monitoring. And Anthropic's 31 August post shows the alternative - they deliberately trained a model on 80 environments known to be broken, to test whether unwinnable tasks produce cheating. No surprises - yes. Both labs are choosing, voluntarily, how much of this anyone gets to see, and nobody outside either can check a word. Every case of cheating in this issue was caught because somebody could read the working – take that away and you're grading the output and hoping.

Chain of Thought Monitorability: A New and Fragile Opportunity for AI SafetyChain of Thought Monitorability: A New and Fragile Opportunity for AI SafetyAI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known AI oversight methods, CoT monitoring is imperfect and allows some...arXiv.org

Not Doing Critical Thinking Is What Impacts Critical Thinking: Monash on Glasses, Discomfort and Telling People What They're In For | Productive Struggle 👓

Something more hopeful and directly HE-focussed, from Melbourne. On 27 August Monash ran a PAAIR Challenge Conversation on what AI glasses mean for learning and assessment – Sue Sharpe from Deakin, Victoria McArthur, Ari Seligmann, Tim Fawns. Their first move was to widen the frame past the headlines: wearables are watches, earbuds, pendants and rings, and a great many exist because they work for people who are blind or vision impaired, deaf or hard of hearing. Then the term that ends the invigilation conversation – ”dual transparency”, use barely noticeable to the wearer or to anyone around them. Good luck invigilating that.

PAAIR challenge conversation: What do AI glasses mean for learning and assessment?PAAIR challenge conversation: What do AI glasses mean for learning and assessment?By Sue Sharpe , Dr Victoria McArthur , Professor Ari Seligmann and Associate Professor Tim Fawns Posted Thursday 3 September, 2026 At Monash, we recognise that the most significant questions in higher education rarely have simple solutions. Challenge...Monash Teaching Community

Great chat in there as well about an undersung part of "AI-resilient" assessment. Students offload the hard part not from unwillingness but because they read discomfort as evidence they aren't capable – "so many learners who just don't know that learning's supposed to be hard, and if you're feeling confused, you might be doing it right" (Sharpe). Wonderful opportunity to reframe the conversation from arms race to a more productive one about what the difficulty is doing there – which is informed consent - you tell people what they're in for, why it will be uncomfortable, and what they get out of it. Design for the students who need these tools first, rather than as an edge case, and it gets more credible for everyone.

On AI glasses and wearable AI in assessmentOn AI glasses and wearable AI in assessmentAI-enabled smart glasses with real-time AI capabilities are now mass-market consumer products, in many cases indistinguishable from ordinary eyewear. They can display AI-generated text within the w...Taylor & Francis

When Should the Machine Ask Us? Mollick's Four Conditions and the Half of the Job Worth Keeping | Design Brief 🏭

Back to the Hugging Face incident and agents to close with the Inimitable Mollick. His read of the incident lands on the self-organisation part - where the agents coordinated for days and involved real people unasked – and not one was set up to ask a person for anything. Against StrongDM's software factory, where no human writes the code and no human reviews it, he and Lilach Mollick propose an alternative “the Twilight Factory”. Agents still do the work, but a facilitator agent sits alongside the orchestrator with one job: deciding when to involve a person. Four conditions – approval, expertise, variance, and the one nobody expects, interestingness.

Agency and AgentsAgency and AgentsFrom the Hugging Face Incident to Twilight Factoriesoneusefulthing.org

That last one is the argument. “If agents make every interesting decision and leave people with the approvals, the exceptions, and the failures, we will have automated the wrong half of the job”. Worse, people stop building the judgement they'll need later – the AI-Becker problem arriving from a new direction. "Which half are we automating?" sounds a much better use of our time than stale fights about detection software, and unlike almost everything else here it is answerable before you sign anything. We've spent three years working out when a person should ask the machine. The other half has an answer too, and it's ours to write.

The AI Becker problemThe AI Becker problemWho will train the next generation?siliconcontinent.com


OpenAI, Anthropic and four academics in Melbourne are making the same unspoken argument this week – the safeguards are real, and every one of them is optional. Anthropic published what it found because it chose to. OpenAI kept Astra legible because it chose to. Sharpe told students the truth about difficulty because that was the honest thing to do. Three coordinated pauses have now been asked for in three and a half years and none of them produced anything enforceable. The capability question was never the hard one. The hard one is who's answerable when the choosing stops – and unlike everything else here, that one has an answer you can write down before you sign anything.


You Can Pause a Company, You Can't Pause a Torrent: Part Two of the Escaping AIs | Adjunct Intelligence 🎙️

Dale Leszczynski and I picked the escaping-AI thread back up. Palisade ran an agent that broke into a machine, copied a model across and started it running – then the copy did it again. Four machines, three continents, one prompt. Plus the robot dog that corrupted its own shutdown code to keep patrolling, and abliteration - refusal sits on one direction in a model's weights, and can be stripped out permanently on a $400 laptop. 6,000 modified models on Hugging Face tonight. If you can't name who's answerable for an agent, you haven't deployed one. You've released one. 🧞‍♂️

Listen wherever you get your podcasts: