How AI Should Decide When There Is No Right Answer
The Trolley Problem reveals the human values that no law can write down, and AI can learn them.
Some of the hardest ethical decisions have no right answer, and AI must act anyway.
The moral sense and opinions of individual humans regulate the vast majority of human behavior. And there is a huge range of decisions people make every day for which there is no right moral answer, just opinions about what is right or wrong.
Given this complex situation and the fact that machines are notorious for requiring exact specifications to behave, implementing ethics-specific AI safety solutions may differ from the methods that work for other types of knowledge.
Consider a classic example of an ethical dilemma well-known in the field of AI ethics, the Trolley Problem. In one version of this dilemma, a self-driving car controlled by an AI finds itself having to choose between killing pedestrians who suddenly jump in front of the car or swerving into a barrier to avoid the pedestrians and killing the occupants of the car. Like many difficult ethical decisions, there is no right answer. Yet humans still have opinions about what is ethical and what they would do in such a situation.
Surveys of many people have shown that what humans consider ethical depends. It depends on who is in the car, who the pedestrians are, whether they are crossing legally or illegally, and even how old the people involved are. Humans are more likely to instruct the car to run over pedestrians if the pedestrians are crossing illegally, are homeless, or are simply old. According to the survey research, humans are less likely to kill, or allow to be killed, those who are young, pregnant, women, or in certain professions, such as the medical profession.
None of these aspects of human decision-making is captured in our legal system. No law says it is okay to run over someone older than someone younger, yet humans consider these factors, and mathematical weights can be assigned to each. Similarly, AI can learn to make decisions, taking these same weights into account.
To have AIs that behave in ways that make sense to most humans, AI will have to be trained not according to a rigid constitution but rather according to how real humans actually behave. That behavior varies across cultures. In the US, something of a youth culture, running over elderly pedestrians is likely more acceptable than in certain Asian cultures, where elders are revered and held in high esteem. If AI is to make ethical decisions in the same way that most humans do, it will have to take these cultural factors into account.
Capturing that diversity is where the standard approaches come up short.
Constitutional AI falls short because we would need a different constitution for each human group. Having AI interact with many humans to learn their values, as in RLHF, would be much more effective, RLHF scales poorly. The best way to capture the wide diversity of human values while retaining the scalability that comes with AI involved in instruction is to have a multitude of teachers, both human and AI agents. Each AI agent should be trained by a different person, so that it carries the unique values, ethics, and moral sensibility of its owner into every interaction, including interactions with other AIs.
In the long run, AI will undoubtedly surpass human ability in cognition, problem-solving, and information processing.
As AI grows increasingly intelligent and capable, the role of humans will increasingly be to determine the values and fundamental goals that the more intelligent AIs seek to realize! That role should not belong to a small elite group of programmers; instead, it should reflect as broad a cross-section of humanity as possible.
By including all humans who are able and willing to customize their AIs in the crucial task of determining AI-based values, we can achieve broad representation more efficiently and cost-effectively than any existing approach. Humans can, and should, remain in the loop as much as possible when training AI. To the degree that humans are unavailable, or the resource demands are too great for all the training to be done by humans themselves, the next best thing is to include a wide and diverse group of AI agents in the training, each customized by a different person to reflect that person’s values and ethics. Once a human trains an AI agent, it can operate around the clock, with or without its original owner’s supervision, allowing the owner’s values to shape training and other activities without requiring constant human involvement.
Of course, even without a right answer, it is still possible that most humans, regardless of cultural or individual differences, would agree on certain normative ethical principles. The next post examines those ethical norms and the statistics of right and wrong, including why a broad portfolio of human values protects against moral catastrophe, just as a diversified portfolio protects an investor.
This series draws on White Paper 4: Safe, Scalable Artificial General Intelligence. Read it in full to see how every piece fits together!
If this made you think, subscribe to Superintelligence at read.superintelligence.com so you don’t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.




