Who Should Teach AI Right from Wrong?
AI learns better values when more people help teach it.
Safe AGI needs many teachers.
In my design described in the SuperIntelligence White Papers, intelligence is the property of a community. The source of safe AGI is a collection of many intelligent agents, human and AI together, and no single monolithic model.
Millions of individuals each customize a personal AI agent, teaching it their knowledge and values, using the methods described earlier in the post below.
Your AI Should Think Like You
What would it mean to have an AI that doesn’t just assist you but actually represents you? When you ask a generic LLM for a recommendation, it draws on the same pool of internet-sourced information it gives everyone. It does not know much about your specific needs, your ethical commitments, or the expertise you’ve developed over the years. A customization mechanism can solve this problem.
Each of those agents carries the moral fingerprint of one person. Today’s leading safety methods fall short of that standard.
RLHF keeps humans involved and cannot cover enough scenarios, while Constitutional AI covers more scenarios by handing the teaching to a single AI and pushing humans aside.
Both rely on too few voices to represent humanity.
Those customized agents, together with the humans who taught them, then form the community that trains new AIs.
When a new AI is being trained, the feedback that shapes it comes from many agents at once, some human and some artificial, in a process I call Reinforcement Learning via Feedback (aka RLF). Because AI agents can give feedback rapidly, scenario coverage can scale in much the same way as Constitutional AI. Because millions of humans stand behind those agents and can participate directly whenever resources allow, the process can incorporate a much broader range of human values.
The risk of relying on one teacher becomes clear when that teacher must interpret general rules in situations their authors never imagined. A single Trainer AI must generalize a short list of written rules to situations its authors never imagined, and it is difficult to know whether it is interpreting those rules appropriately. An AI might learn that preserving the environment is good, observe that humans are harming the environment, and conclude that the best way to protect the environment is to reduce the human population. Although logical, the conclusion is not one that most humans would consider ethical or acceptable. Even when obvious conflicts like that one are explicitly trained out, it is very difficult to anticipate how values will play out across complicated chains of reasoning, actions, and effects.
Of course, humans create unanticipated problems too.
Gasoline-powered cars solved a transportation problem and created pollution problems that were not initially anticipated. However, humans have generally had time to react and adjust when consequences emerged. AI thinks and acts much faster than we do, and a miscalculation could cause severe damage before any human detects it, let alone corrects it.
There is a second problem with AI teaching AI, and it compounds quietly.
In the children’s game of Telephone, a message passes from ear to ear, arriving subtly distorted. A similar distortion can occur when training passes from AI to AI across generations. A human says, “I think XYZ is true.” An AI summarizing that statement records “XYZ is true.” By the third generation, the original doubt may have disappeared, and a machine somewhere is acting on borrowed certainty. The subtleties that come with human involvement are among the things that can be lost as successive generations of AIs process complex, ambiguous data.
Many of us have already lived through a small version of this.
A recommendation algorithm suggests movies, and at first, the suggestions are useful. After a while, we find ourselves wondering why the range of choices has become so narrow. The algorithm picked up on the central tendencies in a few of our frequent choices, ignored the subtleties, and fed us a steady diet of one kind of content. We consumed it, which further convinced the algorithm. The AI initially slightly biased us, then compounded its imperfect understanding, amplifying its own error. With movies, the amplification is annoying. In systems that act on the world, it could be fatal. Without humans in the loop to correct misunderstandings, AI training can drift far off track before anyone notices.
A community of many teachers addresses both problems at once.
A constitution written by an elite few is almost certain to miss the wide diversity of human opinions and cultural norms. At the same time, feedback from millions of agents, each taught by a different person, can reflect a much broader range of human opinions and cultural norms. Human teachers can remain in the loop to catch strange interpretations and lost qualifications before they harden into values. Moreover, humans and AI agents can surface new ethical scenarios as they emerge in real-world problem-solving, enabling the value system to continue learning as the world changes. The values guiding AI can come from a far broader share of the people the system will affect!
The result can be cheaper than RLHF and safer than Constitutional AI, with far more scenarios covered and with humans participating as fully as resources allow. There is also a benefit beyond safety. Broader participation may also increase public acceptance of the AI whose values people helped shape.
Combining values from millions of teachers, however, requires rules for how much each voice counts. The next post examines the two main approaches to pooling what the community knows: one based on feedback and the other on combining the AI agents' weights directly, along with the weighting schemes that determine whether an expert outweighs an everyday contributor and whether a human outweighs an AI.
This series draws on White Paper 4: Safe Scalable AGI. Read it in full to see how every piece fits together!
If this made you think, subscribe to Superintelligence at read.superintelligence.com so you don’t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.





