When we say that AI should have human values, the natural question is whose values, followed by how we resolve conflicts, or whose values win out. These are age-old questions, and the best humanity has come up with, in my view, is the deeply flawed democracy, probably viewed as the worst political system except for all the others. So, similarly, we don’t have perfect answers for AI when it comes to questions of whose values, but we might begin with democratic principles, which means voting.
A simple method of determining consensus values, ethical preferences, and/or the weights or other information reflecting those preferences is to have each intelligent (human and/or AI) entity vote on the values or ethical preferences that should form the basis for the SuperIntelligence’s behavior. Specific scenarios could be presented to each entity, with a range of choices or options for how the SuperIntelligence should behave. Then the entities could vote on the preferred option. Variations are possible, with entities being asked to rate or rank options rather than vote on a single best option. For example, rating or ranking options allow the capture of additional information that would otherwise be lost in simple voting.
One issue that arises when multiple preferences or votes are combined to make a decision is whether each entity’s preference or vote carries equal weight, or whether some carry more weight than others (in a weighted scheme).
Generally, there are always entities that feel they have the most correct, or a more correct, view of values and ethics than others. These entities, or groups of entities, often advocate that their views should count more than the views of others.
I believe that SuperIntelligence should be designed to reflect empirically derived human behavior and that a representative and statistically valid sample should be used.
A preference for using unweighted combinations of preferences (that is, one human, one vote) seems to follow most naturally from these design principles.
However, an important assumption is that the entities doing the voting represent a statistically valid sample of the overall human population. There may be cases where it is known or suspected that the sample of voting entities is not representative; therefore, weighting to correct sampling biases is warranted.
One specific form of weighted voting is to allow voters themselves (human or AI) to assign weights to their votes based on their self-assessment of their qualifications to opine on a particular option, their experience, their level of concern, or other self-determined criteria.
The advantage of such an approach is that it allows the system to incorporate additional information (such as levels of experience or concern) alongside the actual vote. One potential issue with self-weighting is that some aggressive voters may give themselves disproportionate weight on all issues. In contrast, shy or less confident (but potentially more informed) voters might self-censor.
A version of cumulative voting in which voters have the same total number of votes, which they may distribute in different proportions over a range of options and issues, could be used. For example, imagine that a person is given the task of voting on issues of cruelty to animals and racial discrimination. Each person can cast a maximum of ten votes across both issues. A pet owner with strong feelings about animal cruelty might choose to cast all ten of their votes on the cruelty to animals issue. In contrast, a person who has extensive experience with racial discrimination and no experience with pets or animals might weigh the discrimination issue more heavily.
A final dimension, relating to voting and combining input from multiple entities, concerns whether the entities can see and/or discuss other entities’ votes. While anonymous or secret ballots are useful for avoiding peer pressure and related social influences on opinions, there are times when open discussion helps encourage entities to cast more thoughtful votes.
Not all values and ethics for SuperIntelligence need to come from newly created constitutions, voting, or analysis of behavior patterns implicit in datasets. Societies globally have invested substantial time and effort in constructing systems of law, regulation, and ethics from which values and normative standards of behavior can be reverse-engineered. While not every law is just, the overt purpose of laws is to facilitate justice. Therefore, especially within a given society or cultural milieu, that society’s laws provide a good starting point for determining its values and what constitutes ethical behavior. In most cases, legality represents the minimum standard or the limits of what behavior is tolerated by a society.
Further, the distinction between (for example) misdemeanors and felonies helps clarify which behaviors are “more wrong than others.” What is true of laws is also true of religious scriptures and philosophical/ethical texts that attempt to define ethical behavior and principles for community members. Both laws and ethical texts can be used to train AI to provide an initial ethical basis that can be further refined using other datasets and input (via surveys or voting by intelligent entities). Philosophical texts ranging from Aristotle’s Ethics to Kant’s Critique of Pure Reason, as well as political documents such as the UN’s 30 Articles of the Universal Declaration of Human Rights, contain a wealth of ethical information that AI could analyze and extract.
One potential concern with the general approach of giving equal weight to the ethical preferences and values of all humans is that some humans lack the experience and judgment to make “good” ethical decisions. Just as many countries have a minimum age requirement before allowing their citizens to vote, it is possible to design systems that take age and experience into account when determining which ethical preferences and values should carry more weight.
Despite the many possible weighting schemes, my bias is that, in the recommended implementation, aside from perhaps ensuring that humans can demonstrate they understand the ethical scenarios and questions, the weighting scheme for voting on ethical issues should remain as representative and statistically valid as possible. Typically, as a first approximation, one human, one vote, without additional weighting, is a good way to achieve this result. Giving some humans more power than others is a slippery slope, since there are many opinions about how votes could be weighted and many ways individuals or groups could be marginalized if their votes were given less weight.
One approach to enabling more trustworthy and experienced agents (human or AI) to have greater voting power in ethical decision-making is to allow them to delegate it to other trusted agents. Many existing democracies, for example, have humans vote for other humans who then represent them in voting on issues. This procedure is a type of delegation of voting authority. However, in contrast to elections in which one candidate wins and votes on behalf of all constituents, agents could delegate their voting power to a wide range of other agents. One of the advantages of delegation is that it allows people to exercise their free will in choosing who votes on their behalf, while still preserving many of the benefits of having more voting power behind decision-makers with greater experience and expertise in certain areas.
For example, I may choose to delegate my voting power on medical ethics questions to a friend, who has spent his entire career focused on such issues and whom I trust to make more informed decisions in that area than I could. He might delegate his voting authority to me for ethical issues surrounding autonomous AI agents, if he felt I had more expertise in that area. Because each of us has control over whether we delegate authority, we avoid the slippery slope of outsiders determining how much weight our votes carry. At the same time, we can achieve a better outcome by delegating to those we know have more experience in certain areas.
One might imagine that multiple AI agents exist that have proven themselves to vote in ways that align with a human’s views, while also having greater expertise in specific areas. Humans might then choose one or more AI agents to act on their behalf to represent them in certain ethical decisions. Should such delegation occur, in the recommended implementation, the human delegators should still be warned or notified (with the intensity and/or frequency of the warning or notification increasing in proportion to the seriousness of the ethical decision) when the agent makes a decision or votes on behalf of the human who delegated authority. In the recommended implementation, the human delegator should have control over setting the conditions under which they are warned or notified.
Finally, religious, political, or other groups may train and tune AI agents to reflect a specific set of values or ethical preferences. For example, a group of Christian humans, all belonging to the same church or branch of Christianity, might adopt a set of values determined by their church as the default system for each of their individual AI agents. The individual church members might change only those values where their opinions differ from the default view. Alternatively, the individual human might accept the Church’s default values without modification, allowing the Church to “vote” these values with weights proportional to the number of humans who have delegated their voting authority to it.
Voting can be hierarchical, with multiple levels of groups and subgroups. There can be unintended consequences of such an approach. For example, suppose that 3 out of 5 agents in each of 3 groups vote for X and 5 out of 5 agents in two other groups vote for Y.
If all five groups are combined into a higher-order group with 25 votes of aggregate voting power, then, because 3 of the 5 groups voted for X, the entire 25-vote block is cast for X.
But actually, if we count the total number of individual votes (within the smaller groups), we find that only 9 of the 25 votes were cast for X, whereas 16 were cast for Y.
Thus, by using a majority rule combined with hierarchical group voting, it is possible to arrive at a result that actually reflects the opinion of a minority of voters.
While the use of hierarchical group voting can be efficient and desirable in some situations, transparency about how the voting maps to the actual votes of individual agents should be required to ensure that human and AI agents understand the process and agree that it is working as desired.
Most of us are familiar with the situation where YouTube, Netflix, or another video streaming service recommends content based on our stated preferences and profiles. Just as these “recommender algorithms” recommend content, they can also recommend weights on values and other sets of knowledge, preferences, and behavioral profiles, which we could explicitly instruct AI to use in an effort to capture our ethical preferences and values.
I believe that the use of such recommender algorithms should be transparent to humans and require their approval before being used to influence or train AI. However, currently, many such algorithms are used to recommend content in an automated and non-transparent way. Therefore, it is likely that AI would use such algorithms to infer moral preferences without the humans being aware of what is going on. From an efficiency standpoint, such automated use of recommender algorithms would likely be more efficient and, in some cases, more effective than requiring explicit human approval. Thus, designers of such systems must determine not only what is most efficient and effective but also what goals and principles they want the system to reflect.
With respect to automated content recommendation algorithms, a current problem is that they have been explicitly programmed to recommend content that leads to the longest watch times, most engagement, and highest conversion to purchase behavior. For this reason, humans quickly find themselves trapped in “echo chambers” where they are shown more and more content that is similar to other things they have already watched, and for which large amounts of ads can be shown. Society is the loser, and advertisers are the winners, probably not the scenario we want to be repeated with, and amplified by SuperIntelligence.
One method that might help ameliorate this problem is to allow humans to explicitly state their goals regarding the content they see (or, in the case of recommender algorithms, the general values they wish to emphasize). For example, when it comes to content recommended by YouTube, I would like the ability to specify that I want to see a variety of views on a specific topic, and that those views should be representative of the actual views out there rather than what I have already watched. Similarly, when asking for groups or individuals whose values I might want to use as a basis for delegating (some of) my moral authority, I might want to restrict those groups to those that advocate only certain principles. Within that limitation, I may want to evaluate as many different variations as possible so that I can choose from a wide range of options and not just delegate to what an algorithm thinks I like.
Generally, as recommender algorithms incorporate more advanced AI that can reason and respond to requests rather than optimize content based on ad views, the echo chamber problem (and related bias problems) should diminish. In the recommended implementation, SuperIntelligence should avoid echo chambers and seek to gather as much thoughtful, deliberate input as possible from the humans (and their AI agents) who provide moral input. The voting analogy is that society benefits generally when voters are more educated and know more about what they are voting on. Thomas Jefferson’s view, paraphrased as “An educated citizenry is a vital requisite for our survival as a free people,” is applicable here.
The question that all democracies with voting have is, what about the minority?
If the majority always wins, the minority is usually exploited, as we have seen.
So every democracy or democratic system needs mechanisms to protect minorities and ensure their voices are heard.
That is the subject of the next post.
This series draws on White Paper 7: Safe Alignment of SuperIntelligence. Read it in full to see how every piece fits together!
If this made you think, subscribe to Superintelligence so you don’t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.






