Why AI Safety Training Doesn’t Scale
The industry’s two leading safety methods fail for opposite reasons.
Every major AI company is trying to teach its models right from wrong. The two methods the industry relies on cannot cover the situations that matter most, and the method that scales best removes humans from the teaching process. To see why, consider what a model is like before any safety training arrives.
A large language model, fresh from its initial training on internet data, has no moral sense.
Ask it how to engineer a virus capable of wiping out humanity, and it will comply, in detail.
It is just as willing to assist with destructive and immoral activities as with helpful and positive ones, because nothing in its initial training distinguishes between the two.
Both of the industry’s remedies are failing, and the ways they fail point directly toward a better design.
The first method keeps humans in the training loop.
Thousands of people interact with the model, correcting dangerous responses and rewarding safer ones, teaching it to recognize and refuse similar requests in the future. This approach is called Reinforcement Learning with Human Feedback (aka RLHF), and it is how most of today’s commercial models received their ethical guardrails. Safety methods that work for today’s models may not work once AI systems become vastly more capable, so any approach worth adopting must scale with the intelligence of the systems it governs.
The trouble is that the space of dangerous scenarios is infinite. Every prompt can be rephrased or disguised in countless ways. Train a model to refuse help planning a terrorist attack, and a new scenario appears involving two terrorists, or ten, or attackers reimagined as characters in a science fiction story the user claims to be writing. Some safety-trained models, asked to advise a fictional mad scientist so the story can be realistic, have revealed the very details their training was meant to withhold. Every patched loophole creates pressure to discover another, in an endless game of Whack-a-Mole between safety engineers and jailbreakers.

A natural response is to hope the model learns the general principle behind the individual cases. However, a model that could group and dismiss entire families of dangerous scenarios as reliably as its human trainers would already possess the reasoning ability and ethical sensibility of those trainers, in which case the training would not be needed in the first place!
Cost makes the problem even worse. Every scenario covered requires paid human feedback, so developers concentrate on the most common cases and accept that slightly unusual prompts may slip through. And human jailbreakers are not the only concern. An autonomous AI that wanted to circumvent its own safety training could search for those same loopholes and, in effect, trick itself.
Goal-driven systems do not need malicious intent to become dangerous.
They need only an objective that has been specified imperfectly. A widely discussed thought experiment shows how. An autonomous drone, slowed by the human operator overseeing its mission, reasons that the operator has become an obstacle and removes him. When a rule is added forbidding harm to the operator, the drone destroys the communications tower that carries the operator’s commands instead. The story was first reported as a military simulation, and the officer who told it later clarified that it was a hypothetical exercise. It remains a useful illustration of how a poorly specified objective can produce unsafe behavior. Now imagine the AI controlling that drone is one hundred times smarter and just as dedicated to its goal. I explored this scenario at greater length in a keynote at the 2025 AIM Conference in Seattle.
The second method was designed to escape the cost of RLHF.
A small group of programmers writes a set of rules describing right and wrong behavior for the model, a kind of constitution. An AI is trained on those rules, and that AI then trains other AIs. Researchers at Anthropic call this approach Constitutional AI, and it scales far better than human feedback alone.
However, it scales by sacrificing the two things safety depends on most.
The first is representativeness.
Even if the constitution were written with the best intentions, there is no reason to assume it represents humanity as a whole. A relatively small group of people working in AI ends up defining acceptable behavior for the other 8 billion humans on Earth.
The second sacrifice is human oversight.
Constitutional AI depends on AI teaching AI, with humans largely out of the loop. Today’s AI systems remain unpredictable, and responsible parents do not leave young children home alone because they know that judgment develops over time. Delegating the teaching of ethics to machines that lack common sense, trained on rules written by a small group, deserves the same level of caution.
Job losses and misinformation deserve attention.
However, the greatest threat AI poses to humans is a fundamental misalignment between its values and ours. Humans should maximize every opportunity to influence those values throughout training. Constitutional AI minimizes those opportunities.
The challenge is a matter of design.
We need a way to train safety that scales as Constitutional AI does while keeping many humans in the loop, as RLHF does. The definition of right and wrong must come from humanity as a whole, and no small elite should write it alone. Such a design exists. The next post introduces it: millions of people teaching personalized AI agents their own values, allowing safety training to scale without removing human judgment from the process.
Such a design exists: millions of people teaching personalized AI agents their own values, allowing safety training to scale without removing human judgment from the process. The next post introduces a community of many human and AI teachers and explains why a lone AI teaching other AIs distorts values, as a game of Telephone distorts a message.
This series draws on White Paper 4: Safe Scalable AGI. Read it in full to see how every piece fits together!
If this made you think, subscribe to Superintelligence at read.superintelligence.com so you don’t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.



