Whether we like it or not, looking at what humans actually do (including what they say and write, since verbal communication is actually a form of behavior) has been. It is likely to remain the primary way that AI learns what humans believe is right or wrong. If we want to understand human ethics as it is actually practiced, we must examine human behavior and adopt an empirical approach.
Even when brilliant philosophers like Immanuel Kant write immense treatises such as the Critique of Pure Reason, their ethical systems are inevitably based on assumptions. If one does not agree with these assumptions, the entire intellectual edifice crumbles. Thus, no system of philosophy has developed a universal ethics that all, or even most, people would agree with. Instead, the philosophers seemed to offer a smorgasbord of interesting ideas. This observation leads to the conclusion that to understand human ethics as actually practiced, it is more productive to observe what humans actually do rather than what they, or philosophers, think they should do.
“Practice what you preach” is a well-known injunction. Also, the phenomenon of children paying more attention to what their parents do than to what they say is commonplace. It seems that, for all our high ideals and philosophizing, we must look at human behavior.
When ethical dilemmas are posed, such as the Trolley Problem (in which one must choose between running over people in a crosswalk or crashing the car to avoid running them over but thereby harming the occupants of the car), human beings don’t turn to Kant’s Categorical Imperative. Instead, they act and react. Similarly, human character is revealed not in easy choices but in difficult circumstances, often involving temptation, fear, greed, pain, or other factors. Thus, because AI learns by observation of what humans do (just as children do), a fundamental principle is that whatever ethics or values we wish to communicate will and (arguably) should be based on empirical data reflecting what humans actually do.
Some may worry that with all the terrible things humans have been known to do, we are setting a terrible example for AI that will result in human destruction. However, these people are forgetting all the wonderful, positive, and loving things that humans do as well. In fact, if the preponderance of human behavior were negative, homicidal, or suicidal, humans would not have survived as long as we have. If human nature were fundamentally evil, and we have had the technology to make the species extinct for many years now, wouldn’t we most likely already be extinct? The facts that we are concerned not so much about deliberate use of nuclear weapons as we are about nuclear accidents, and that we agonize and become deeply depressed about holocausts and genocide, are indications that humans, at their core and in the main, are not a homicidal or suicidal species.
As a Jew, I regard the Nazi Holocaust, which killed six million Jews, as one of the most horrific events in the last hundred years. Yet even this tragedy, which most consider among the worst things that humans have ever done to each other, killed about a quarter of 1% of the world’s population. If humans, by nature, were essentially genocidal, then most of the world would not have the huge collective guilt and horror that they experience when contemplating this event. Moreover, if humans were genocidal by nature, many other events, killing a substantial portion of the human population, would have occurred.
According to the Pan American Health Organization, the United States lost more people to the Spanish flu in 1918 than in World War I, World War II, the Korean War, and the Vietnam War combined. However, there were few, if any, protests against the flu and the conditions allowing it to spread, whereas each of the aforementioned wars generated many protests and political debates. The fact that the reaction against human-to-human violence is so much greater than the reaction against a flu bug suggests there is a widespread moral concern over murder and genocide that is not just related to the number of human deaths. Our reactions to negative events are very important in setting an example for AI that learns, like a child, from both positive and negative examples.
An AI, devoid of emotion but expert at learning patterns in the data of human behavior and speech, cannot help but conclude that, despite the occasional bad behavior of individuals like Hitler or school shooters, the vast majority of humans may trash-talk or cut each other off in traffic, but rarely kill each other deliberately. Moreover, whenever such human-against-human violence occurs, there is almost always a universal outcry from other humans against the perpetrators of the violence. Even in wars, while the warring parties attempt to justify their actions, the rest of humanity on Earth (who do not have a vested interest in the war) condemns the behavior and tries to mediate the dispute and end the violence. Witness the December 2023 vote in the UN General Assembly for a cease-fire in Gaza during the Israel-Hamas war. One hundred fifty-three nations voted for the cease-fire, while only 10 nations (with a vested interest) voted against it. This pattern has recurred so many times in human history, and the record of it is so unequivocal that AI analysis of actual human behavior cannot help but infer that, while humans may occasionally engage in violence, the preferred and normal behavior is to engage with one another in peaceful and mutually beneficial ways.
Finally, we note that humans, when they are nasty, tend to be far nastier in their speech than in their actions. Who among us has not said things in anger that we regret? Yet most of us can refrain from (at least the worst) actions based on these words. AI must distinguish between actions and words. This is a final and important reason why physical behavior should be given more weight than words when inferring ethics and values. Common wisdom holds that “talk is cheap,” “actions speak louder than words,” and (as children say on the school playground) “sticks and stones can break my bones, but words can never hurt me.”
If one agrees that human behavior is mostly “good,” or at least mostly not genocidal or suicidal, there remains the major problem of AI potentially being misled by a non-representative sample of the data. If a model, for example, were trained exclusively on datasets containing information about war, genocide, and horrible things humans have done to each other, then, because of the biased sample, the AI might generalize to the wrong set of values.
Less extreme but also problematic is the situation in which the training data comes from only one race, gender, age group, culture, or geography. As many researchers have noted, such biased datasets can lead to AI systems that are prejudiced in ways that most humans would consider morally “wrong.”
There is a temptation to go to the other extreme and carefully select the datasets used to train AI so that they only contain humanity’s noblest actions and words reflecting our “highest ideals.” But who is to choose what is noble, what specifically goes into the data, and what is excluded? Further, what happens if ideas of what is noble change over time, or from culture to culture? Would not such an approach lead to battles over which human ideas to include and whose ideology should prevail? The potential for slippery slopes is profound.
I argue that a representative and statistically valid sample of actual human behavior should be used for training AI systems. This means including all human behavior, warts and all, but in a way where the included data is proportional to the actual occurrence in the population.
Those who fear the inclusion of negative data in the training set may misunderstand my proposal as calling for the inclusion of a representative sample of what the news reports. That is not what I mean. The news is not a representative and statistically valid sample of human behavior. Rather, the news is heavily biased towards negative, shocking, and unusually bad human behavior, which tends to grab human attention and sell ads.
In the near term, when humans are responsible for selecting and filtering the datasets used to train AI, great care should be taken to ensure that the behavioral data is, in fact, representative and statistically valid. Such data would include many instances of ordinary, peaceful interactions, with only a tiny fraction of aberrant and violent behavior.
In the long run, AI systems will be more than capable of determining representative and statistically valid samples of human behavior data for themselves and, lacking the same emotional triggers, would likely be far better than humans at constructing an accurate picture of human nature as it is, as opposed to how the press portrays it for advertising purposes or how special interests attempt to manipulate it for their own purposes.
In any event, as definitions of what is moral or noble change over time, an empirical approach (using statistical methods) to assessing human behavior (including verbal communication) seems the most practical and sustainable.
If the answer is to design smarter AI that is safe and aligned with human values, then how do we do it? The next post lays out ten principles for doing it and why it is the smarter design that leads to safety.
This series draws on White Paper 7: Safe Alignment of SuperIntelligence. Read it in full to see how every piece fits together!
If this made you think, subscribe to Superintelligence so you don’t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.





