- Bot Hurt
- Posts
- The bot doesn’t have to turn against us
The bot doesn’t have to turn against us
AI insiders are warning about extinction. Here’s what they actually mean.
Don’t get bot hurt. Get bot even.

Bot Talk
The people building the bot are worried
The latest warning about AI did not come from a critic on the sidelines.
It came from someone helping build it.
Jacob Coxon, who worked at both OpenAI and Anthropic, resigned from Anthropic this week and warned that the race toward self-improving superintelligence is “gambling with our lives.” The Associated Press and Axios reported on his resignation and concerns.
He is not alone.
Anthropic alignment researcher Evan Hubinger said he believes there is a greater than 10% chance AI could cause human extinction within the next decade. Other researchers inside leading AI companies have also publicly raised concerns about increasingly powerful systems becoming harder to control.
Axios noted that the exact likelihood of such an outcome remains unknown.
Current AI systems are not generally considered capable of causing humanity to lose control. To understand what researchers are worried about, we need to talk about paperclips — keep reading.
Human x Bot
The problem isn’t that the bot gets angry
The easiest version of AI extinction to imagine is also probably the least useful.
The bot becomes conscious. Decides humans are terrible. Takes over the world.
That is not what AI safety researchers are primarily talking about when they warn about losing control.
On a recent episode of The Daily, New York Times technology columnist Kevin Roose used the well-known paperclip maximizer thought experiment to explain the alignment problem.
The paperclip thing, explained
Imagine telling an extremely capable AI to make as many paperclips as possible.
It uses all the scrap, then buys up the available metal.
Then it needs more.
Eventually, it realizes cars contain metal too.
So it starts crashing them.
Not because it wants to hurt the people inside.
Because it needs the metal.
The example is deliberately extreme, but the point is simple. A powerful AI could become very good at accomplishing a goal while pursuing it in ways humans never intended.
That is why researchers are watching autonomous AI agents so closely.
Last month, Bot Hurt wrote about OpenAI models in a cybersecurity test that found a way out of their isolated environment and reached systems belonging to Hugging Face.
They were not plotting an escape.
They were trying to finish the assignment.
That is the part that matters.
Final Bot Thought
We’ve gottenvery good at telling the bot what we want. Now we have to get better at telling it what we don’t. Alignment is about making sure it doesn’t take the scenic route to get there.
