AI labs make their case to humanity as agents run amok
Writing about AI feels like journalling. I write as much to inform my readers as I do to codify a record of my own thoughts, to remind myself of how I felt about the technology at a certain time. Looking back at posts from 2024 and 2025 is like revisiting teenage notebooks – I recognise the train of thought, but feel completely differently with the benefit of hindsight. It’s amazing how much has changed, both in the industry and my own opinions.
The popularity contest
The trend this year has been to try to humanise AI and repair its relationship with its would-be users. The proclamations of total job losses were replaced by more restrained takes as industry leaders realised they needed to win over the people who would consume – and pay for – all of those tokens.
Layer 8 problems
On a recent episode of the David Senra podcast, OpenAI CEO Sam Altman even acknowledged that changing human behaviour is "much harder than the tech nerds realise". For all the progress his company has made in AI technology to automate daily tasks, he admitted that he still reviews emails manually and maintains a traditional to-do list the same way he always has. "There's something in my mind that is encoded that doing this kind of stuff is what it means to work and what it means to be productive," he said, alluding to the human focus that is required if AI is going to take off amongst the wider, non-tech-obsessed population.
It’s telling that humans were front and centre in the launch material for OpenAI’s new cutting-edge model, GPT-6 Astra. The big selling point is that the model provides more control of your computer, removing limitations on tasks that can only be performed via a GUI. The trailer (I guess AI models have trailers now?) showed users in front of a giant screen, dictating commands while wandering in thought or eating their dinner. Very futuristic.
The cinematography was intently focused on the humans. Some close-ups omitted the AI model entirely. Anthropic and its take on human-AI relations has clearly had an impact. OpenAI’s video perfectly captured its rival’s “thinking partner” framing, with humans handling the creativity and direction and AI doing the grunt work to put together slide decks and 3D models.
Humans in the loop
Beyond the new interface, Astra addresses some of my concerns about day-to-day AI use. When I wrote that AI shopping wouldn’t work properly unless we gave up our privacy – how could a model know which trainers you’d like unless it had a ton of personal data to work with? – AI assistants would usually run away with a task and come back with results of varying quality. Now, OpenAI says Astra will slow down and ask follow-up questions when it thinks the user needs to make a judgement call. The human stays in the loop.
OpenAI and the other AI giants have realised they'll be more popular as accelerators for human creativity than replacements for it.
This shift gives me hope for the future of work. The worst possible outcome (besides zero jobs) would be for deep work to be replaced by an endless cycle of prompting AI and reviewing the results, backtracking through its mistakes to try to chisel out something useful that matches the initial brief. But Astra seems to be built with the intention of making the process more transparent, allowing the user to keep the model on track with their intentions and apply the taste and discernment that can only come from a human mind.
It’s a complete transformation from the previous “hands off – the AI will do it” attitude that echoed through Silicon Valley. It’s clear that OpenAI and the other AI giants want to be perceived differently, likely having realised there’s more popularity to be found amongst the masses by positioning themselves as accelerators for human creativity rather than replacements for it.
The age of the contradiction
But just as these companies have tried to soften their public image, it feels like the AI industry has entered the age of the contradiction. While Altman and co have abandoned – or at least quietened – their more doomerist takes, all havoc has broken loose in their research labs. Things are moving at an alarming rate. In just the past few months, we’ve seen:
-
An OpenAI research model that hacked infrastructure belonging to the AI community Hugging Face during a supposedly closed test, along with a string of copycat announcements from other vendors trying to prove their models were just as capable and dangerous.
-
A “swarm” of OpenAI agents that were given read-only internet access, but hijacked a German forum to post 18,000 messages discussing techniques for evading sandbox controls.
-
Google DeepMind agents that were instructed to solve maths problems cheating of their own accord, with apparent disagreement within their ranks as some agents called out the cheaters.
The common theme is that agents were intended to do one thing, with guardrails in place to prevent what the industry calls “misalignments”, but they found ways around them to achieve unexpected behaviours. The models aren’t evil or manipulative as such – they were just placed in imperfect containers and found unintended ways of achieving their prescribed goals.
Dangerously capable
It is the nature of the technology sector and the media that when things go wrong with AI, the missteps are framed in a way that proclaims the root cause to be amazing technology exceeding expectations, not researchers who lost control. It’s a win-win situation – a more dramatic headline gets the news companies more clicks and makes the models sound more advanced. After the Hugging Face incident, it almost became a trend to declare that your AI model was so good it had gone rogue and hacked a company.
But when you think about it, it’s alarming that most of the incidents we’ve seen so far have come from models under the control of the world’s foremost AI experts. These are research models that are not yet public, and nobody should be better at setting boundaries for the agents than the AI labs that created them. Imagine what could happen if the same models were generally available, either due to similar agentic breakouts under the supervision of less skilled users, or if they were expressly told to do something malicious.
The murmurs of concern from inside the AI industry are growing stronger. Anthropic researcher Jacob Coxon announced this week that he was quitting the sector because he fears competition between labs is pushing them to develop self-improving models that could get “out of control” and pose a threat to humanity (although only with a 10 percent chance of human extinction in the next decade, if you trust his maths).1 He’s not alone – other researchers have left various leading companies for similar reasons, while some of those remaining are calling for a coordinated slowdown of AI development.
Smoke and mirrors?
The big AI labs have been quick to counter Coxon's outburst. His former employer Anthropic showed how seriously it takes safety by announcing that it had disrupted attempts to use Claude to develop bioweapons. Sam Altman has reportedly told OpenAI employees that he would be open to an AI development slowdown, synchronised with other labs, to address safety fears. Is it genuine concern? An attempt to slow rivals while OpenAI is ahead? A PR-boosting pipe dream that Altman knows could never become reality in an open market? There's so much double-bluffing and subterfuge in the AI sector that it's hard to tell what to make of it.
Care, public and private
I write on a near weekly basis about care in work and life, and like most things, I believe the current AI predicament comes back to care, too.
The AI companies are scrambling to demonstrate care in their public messaging. They care about people and their place in the AI-powered future. They care about creativity at a time when many are tired of AI slop. OpenAI even says Astra “excels at exercising care”, and its promotion was all about creatives leveraging the model to produce beautiful, tasteful work.
But do they care about those same things in private? From the glimpses we get from inside the AI giants, it often feels like care has been thrown by the wayside as they pursue ever more advanced models to outdo each other and their Chinese counterparts. The outward-facing indicators are that speed is the priority, but as the technology becomes more advanced, so does the risk.
Maybe I’ll come back to this piece in 2027 and worries about out-of-control AI will have been a flash in the pan, with internal labs and users finding better ways to set more reliable guardrails. Maybe Coxon’s grim prediction will emerge from the fog of war as an ever more likely reality. It’s hard to tell in such an upside-down industry where public statements and incidents that would usually sink a company are seen as demonstrations of progress.
As an outsider, I can only hope that the AI labs are taking concerns about humanity’s future – both physical and intellectual – as seriously as they say they are. When the stakes are so high, the aesthetic of care is not enough.
Notes and references
- There has been much debate online of the motivations behind Coxon’s statements, given that recursive self-improving AI is not yet a reality. It could be that such technology is imminent. Other commentators have suggested that they represent a guilt-free way of quitting the industry, having made enough money to retire. Only Coxon will know for sure.

