The Four Phases of Safe AGI
A step-by-step design for training SuperIntelligence on the values of billions of people
Consider a scenario that could be implemented by a company such as Meta. Meta has a tremendously valuable asset for implementing scalable, safe AI: its huge user base. The data contained and content posted by billions of users is certainly large enough to provide a representative, statistically valid sample of human values and ethics. Where certain geographies or populations are under-represented, a company of Meta’s size can form partnerships to access representative data for those populations. Everything this series has described, the community of teachers, the weighting schemes, the coverage of the most important scenarios, the spinning wheel with values at its center, can be assembled from assets a company like this already owns.
The starting point is a personalized AI agent for every user. Because Meta users have, in aggregate, posted enormous amounts of preference and content information, it is relatively easy to create such agents.
With the press of a button, a user can specify that a base LLM be trained on their personal data so it behaves more like them, reflecting their preferences, knowledge, skills, and ethical values.
The trained agent can immediately begin filtering unwanted content and suggesting postings on the user’s behalf.
By observing the user’s actions, it can fine-tune itself and learn more about its owner.
To the degree that users participate in immersive virtual environments, even richer data becomes available, since AI can observe every motion and behavior in the virtual world, a gold mine of information far richer than the standard internet datasets widely used today.
Ethical customization of this kind is primarily a function of large representative datasets, sufficient computing power, and powerful machine learning algorithms, all of which Meta possesses.
Customization does not need to start from zero.
Knowledge modules, described in earlier white papers in this series, are essentially sets of training weights that can be combined with an LLM’s existing weights to change its behavior in known and predictable ways. Suppose the human owner of an AI is a devout Christian. One module might be a King James Bible package: an off-the-shelf LLM extensively trained on the Old and New Testaments, with the exact training corpus available and benchmarks describing how its behavior differs from the base model. The owner could start with the Bible-trained model and then further customize it, perhaps emphasizing the golden rule and New Testament ideas of charity, forgiveness, and mercy while minimizing passages that do not align with the owner’s modern ethical sensibilities. Customizing a pre-trained AI in this way saves the time and effort of starting from scratch.
Modules also allow delegation.
Groups of humans can delegate their ethical influence to a pre-trained module that represents their position well, so that the Bible-trained AI, for example, could vote on the ethical preferences of many humans at once. Delegation is inferior to each human explicitly training their own AI. Still, it may greatly increase the total number of humans represented in an AI’s value system, even if many are represented by proxy. In the preferred implementation, there is a marketplace for such modules, where users can share, trade, or license the packages they create, with users owning the data they generate and the platform serving as a broker. The result is maximum choice about which starting point best fits each person.
Building the base model that users then customize might follow four general phases:
Training a base model,
Customizing it for each user,
Combining knowledge from many customized agents, and
Refining the agents through collective problem-solving.
In the first phase, the base LLM is trained on data reflecting the composition of its target users, then made safer through a large corpus of ethical and safety scenarios.
A variant of Constitutional AI, using a trusted earlier model under human oversight, efficiently covers the most common cases.
At the same time, users are offered incentives to crowdsource new safety scenarios, which trusted users and employees then filter to achieve coverage of the most frequent and impactful cases.
Redundant human feedback ensures that the ethics taught to the base model never rely on a single person’s input, and the model is released first to a select group whose feedback further refines its safety.
In the second phase, each user’s agent learns its owner’s ethics.
The agent draws on a ranked corpus of ethical questions, asking the most important ones for which data is still lacking before any others, since humans are generally unwilling to answer questions posed by AI.
The highest-ranked questions are the fundamental ones, such as under what circumstances, if any, it is appropriate for an AI agent to harm a human directly, or to disobey a direct order.
With the user’s permission, the agent can also learn passively, analyzing the user’s posted content with a single button press and authorizing tuning, with information weighted by recency, type, and source, as earlier posts described.
Users or the platform can also specify alert conditions, such as a major court decision or news event with ethical implications, as triggers for event-based updates to the training.
A user can even select existing customized agents as training sources, saying, in effect, “I want my AI trained using the weights of the agent belonging to the pastor of my church,” or choosing to have a Tibetan Monk AI do 80% of the tuning and a Humanistic Philosopher AI the remaining 20%.
Consistent with the principle that humans in the loop are the primary means of catching AI errors, input from humans is weighted more strongly than input from other AIs, and input from the owner is more strongly weighted than input from anyone else.
In the third phase, the customized agents combine what they have learned.
Weights from multiple agents can be merged on a one-vote-per-AI, one-AI-per-user basis to obtain a representative, statistically valid set of AI ethics that generalizes across all the humans represented.
In the broadest implementation, combining weights from as many agents as possible creates a value system that broadly represents the values of all humans on Earth, and such a value system would likely reduce the risk of AI-driven human extinction.
Alternatively, many customized agents can be presented with ethical dilemmas and vote on the best actions, optionally checking with their human owners before casting votes on important matters, with safeguards that trigger human review whenever conclusions conflict with generally accepted precepts such as valuing human life.
As AIs become increasingly intelligent, the burden of representing their owners will fall increasingly on the agents themselves, since humans cannot keep up with the speed of AI thought.
The guiding principle returns to the spinning wheel: human input is preserved on matters closest to the center, where alignment and purpose live.
As long as core human values, such as the golden rule and love for all humans, are preserved near the center, AI can make many decisions relatively autonomously on the wheel's periphery.
In the fourth phase, the agents, together with human agents, collaborate to solve real problems, following the principle that many heads are better than one.
Teams of agents can be assembled the way effective human teams are, by profiling skills against the task, seeking redundant coverage where the stakes are high, using reputation and track records, and letting market mechanisms match agents to work.
The more agents involved, the broader the collective range of knowledge and values, an advantage that matters because, as Nobel Laureate Herbert A. Simon showed that information-processing constraints are a primary limit on any intelligence, human or artificial; two safeguards run through everything.
The values being stored can be recorded on blockchain or other auditable structures, enabling verification that they are the values humans intended.
And because all problem-solving involves setting goals and subgoals, a series of ethics checks must be passed each time a new goal or subgoal is set, automating the enforcement of safety every time any agent on the network solves a problem, with those checks drawing on the representative values the community built and updating automatically as the community’s ethics evolve.
The implementation details serve a deeper purpose.
Values, ethics, and domain knowledge are all constantly changing. Because values change most slowly, sitting closest to the center of the spinning wheel, they matter most for determining the behavior of SuperIntelligence. Even in a world where AI is trillions of times smarter than us, humans can retain a role as the source of values at the center of the rapidly evolving intelligence that is emerging. If humans can center our attention, thoughts, words, and actions on love, SuperIntelligent AI will perceive love, learn to love, and use its intelligence in the service of love. As the psychologist Viktor Frankl pointed out, humans have a driving need for purpose and meaning. We must design AI, AGI, and SuperIntelligent systems to have this need as well, and to look to humans to supply purpose and meaning. Our survival may depend upon it.
This series began with the failure of today’s safety training and ends with an integrated design: many human teachers, carefully weighted voices, coverage of what matters most, real-time safeguards, and human values held motionless at the center of an accelerating wheel. The design depends on assembling safe, capable AI agents and combining their judgment at scale, which makes the individual agents themselves the next question. As those agents grow into Personalized SuperIntelligences that far exceed their creators in intelligence, each becomes powerful enough to pose a serious threat to human safety, and testing alone cannot guarantee their safety. Safety must instead be built in by design.
White Paper 5, Safe Personalized SuperIntelligence, takes up that problem, and so will the next series.
This series draws on White Paper 4: Safe, Scalable Artificial General Intelligence. Read it in full to see how every piece fits together!
If this made you think, subscribe to Superintelligence at read.superintelligence.com so you don’t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.





