Why Training AI Is Like Testing the Brakes
AI trained by millions of people learns the situations that matter most.
If software were a car, we could live with an interior light failing more easily than we could live with the brakes failing. If we had to choose between testing interior lights and the brakes due to limited resources, we would prioritize testing the brakes. AI safety faces a similar problem. There are more dangerous situations than any training system could anticipate individually, so the practical goal is to cover the situations that arise most often and the ones where failure would cause the most harm.
The design described in these white papers meets that goal with a large and diverse community of teachers. Millions of people each customize a personal AI agent, teaching it their knowledge and values. I call these agents Advanced Autonomous Artificial Intelligences, or AAAIs. Since people live in different circumstances and train their AAAIs based on those experiences, the community should collectively cover a broad range of ethical situations.
Software engineers call this kind of problem path coverage, and at some level, the problem of training AI safely resembles it.
AI must be trained on enough representative situations involving dangerous or ethical decision-making so that its behavior becomes more reliable and trustworthy in those situations. Enough of the situations (paths) must be covered in the training. Current AI systems can hallucinate and behave unpredictably. This design aims to make AI substantially more predictable and trustworthy, especially in safety- and ethics-related contexts.
Generally, when testing software, human developers create test cases to cover the use cases most likely to arise.
Since it is impossible to test every possible use of complex software, developers determine which use cases are most common and which have the highest impact if things go wrong. More dangerous scenarios get more testing than benign scenarios. Frequency matters too. If one interior light is used ten times as often as another, a failure in the more frequently used light would affect people ten times as often, so it deserves more of the limited testing. This logic applies when training AI, whether via RLHF, other AIs, or, as this design suggests, a combination of humans and many customized AI agents.
Many representative humans customize AAAIs, and those humans and agents help train new AIs.
If the participating group is sufficiently broad and properly weighted, situations that occur frequently should also appear frequently in the training. In effect, the system samples both human values and the situations in which those values must be applied. The larger the sample, the more certain we can be that the most frequent cases have been addressed in ways aligned with the human population’s values.
Addressing some dangerous cases and difficult ethical decisions is more challenging.
That’s because dangerous situations and tough ethical decisions are often relatively rare. In this case, the approach is to ask humans and AI agents to think of as many dangerous scenarios as possible. The total pool can then be allocated among the human and AI samples so that identified high-impact scenarios receive input from enough different agents to provide a representative range of judgments. Input can then be directed toward people with relevant experience.
Suppose an AI must decide which patients receive medical attention during triage. Emergency room doctors and paramedics, who are used to making triage decisions, may recognize that devoting scarce resources to a patient with almost no chance of survival could reduce the chances of saving other patients. A well-meaning person without triage experience may understandably find that choice harder to make. For this difficult ethical situation, specialized knowledge is an advantage, and we might prefer to let the medically experienced professionals teach the AI. People also tend to identify dilemmas from the worlds they know. Emergency clinicians will think of triage, while an HR professional might contribute scenarios involving hiring, promotion, and fairness. A large and varied population, therefore, generates a broader set of ethical cases than a small, centrally selected group.
A broad community does two things at once: it generates a wider range of ethical situations and connects them with people who understand them. In ordinary life, people gain the most experience with choices they face repeatedly and give special attention to decisions with serious consequences. A community of teachers can create the same pattern in AI training.
Professional human trainers can use RLHF to address gaps in handling important but infrequent ethical dilemmas. This gives the student AI broader and more representative ethical training. Of course, no amount of advance training can anticipate everything.
The next post examines how AI can detect danger in real time, flag situations it cannot judge, and pause for human guidance.
This series draws on White Paper 4: Safe, Scalable Artificial General Intelligence. Read it in full to see how every piece fits together!
If this made you think, subscribe to Superintelligence at read.superintelligence.com so you don’t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.




