AI and the Statistics of Right and Wrong
Diversified human values protect AI safety the way a diversified portfolio protects an investor.
Although it is undesirable to have the safety and ethics of AI driven by a constitution written by a small, elite group, it is still possible that most humans, regardless of cultural or individual differences, would agree on certain normative ethical principles.
Examples of these might include variants of:
The Golden Rule (Do not do to others that which you would not like done to you)
First, do no harm (as reflected in the Hippocratic Oath taken by physicians)
Do not kill (unnecessarily or except in specific, exceptional circumstances)
Preserve individual freedom (unless it limits the freedom of others)
Of course, almost as soon as you read these principles, exceptions spring to mind. Do not kill, but what about self-defense or war? Preserve individual freedom, but what are the limits, and when does it impinge on others? Details and nuance matter, even when applying principles that most humans would broadly embrace.
However, by starting with general normative ethics widely accepted by a large, diverse, and representative group of humans, it is possible to refine these principles and determine when and how they apply in detailed circumstances much more efficiently than if no starting principles existed at all. General ethical norms are a point of departure that can help AI achieve realistic, nuanced ethics and behavior aligned with what most humans believe is good and aspire to.
Once we have admitted the potential usefulness of ethical norms as a starting point for further refinement, the door is open to group and planetary norms.
There is a continuum: at one end are highly individualized AIs trained to think and act like a particular person; farther along are AIs trained to reflect the values of specific groups; and at the far end are AIs guided by ethical norms shared across many groups. Ethical norms at each point on the continuum can serve as a starting point for training AI to exhibit ethical behavior. The idea that one set of norms or one constitution should power all of AI is likely unrealistic and far too brittle to work in the real world.
If it were possible, then the many differing viewpoints espoused by religious, political, and cultural groups would long ago have merged into a consensus. The diversity in human ethical norms is not a bug; it is a feature. We should not expect AI to achieve consensus and maintain human alignment if humans themselves cannot, especially if there is debate over whether such a consensus is even desirable.
Humans also often enter into ethical, implicit, or explicit social contracts when they join a group or participate in society.
Members of a particular religion largely agree with a set of rules and ethical precepts espoused by that religion, often enshrined in holy books. Similarly, Confucianism in China, the ideals reflected in the Declaration of Independence and Constitution in the USA, and liberal or conservative ideologies for various political groups all contain normative prescriptions for human behavior. By being a citizen of, or simply living in, a particular country, humans are explicitly subject to the laws of that country, including laws that explicitly specify what criminal (aka wrong) behavior is. Thus, for ethical problems, the solution sometimes depends on what social or ethical contract humans have made with the group or culture in which they find themselves. Such contracts can be useful for simplifying AI training, since a starting point can be the laws of a particular country or the implicit or explicit rules of a particular group.
It has been said that democracy is a bad political system, but that all the others are worse.
Most humans would agree that the most important concern regarding AI is the existential threat it currently poses to the majority. That is, AI could wipe humans out. If that happens, it doesn’t matter what form of government or religion you prefer. We’d all be dead, and the point would be moot. So we should be asking not which religion or form of government is best, but which principles are most likely to lead to humanity’s survival.
Democratically representing the opinions of most people is rarely optimal, but generally achieves an acceptable outcome.
Collective intelligence (the idea that two heads are better than one) is responsible for the vast majority of human progress, culture, and technology. But when it comes to the subjective area of ethics and values, where there is no objectively correct answer, just human opinions, democracy, or a collective intelligence approach, if you prefer, really shines. One benefit of a democratic and representative set of human values is that it tends to mitigate extreme positions, which are likely to pose the greatest risk to human survival. There is a beneficial diversification effect regarding values.
Just as diversification in an asset portfolio reduces volatility and risk, so too does a diversity of human opinions and judgments tend to stabilize the overall portfolio of values. In an asset portfolio, a diversified portfolio always returns less than if you were to concentrate all the investment on the top winners. The problem is that no one knows who the winners will be with any certainty. That is why the diversified approach of just buying the index tends to outperform more than 80 percent of all portfolio managers who try to beat the market.
Similarly, there are philosopher kings or religious saints who can make laws and ethical rules that, for a time, are far superior to the collective judgment and behavior of the masses. But what happens when the superior king or saint is gone? Then a power-hungry dictator might arise whose reign is far worse. The more stable approach (less likely to be really great, but also less likely to be really terrible) is to follow the values of a large representative population of humans. All these humans want to survive. Most want good things for themselves and their fellow humans. Few want to destroy the environment or the planet. While the collective values are imperfect, they are usually not malevolent. Importantly, they are based on human hearts!
Putting aside the practical benefits of a diversified, representative portfolio approach to human values, a representative sample is also a scientifically valid way to accurately answer a question with no logical answer: what is right and what is wrong according to humans. A fast computer could answer a math problem faster than a million humans, but when it comes to the subjective determination of what is wrong and what is right, calculation speed is useless. If we want to know what human values are, there is no substitute for asking them and watching their behavior. The more humans we ask and watch, the more representative the values may be.
Of course, there are potentially an infinite number of dangerous situations we need to train AI to handle safely, and training resources are limited. The next post examines path coverage, or how to decide which situations matter most, and why, if software were a car, we would test the brakes before the interior lights.
This series draws on White Paper 4: Safe, Scalable Artificial General Intelligence. Read it in full to see how every piece fits together!
If this made you think, subscribe to Superintelligence at read.superintelligence.com so you don’t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.




