In addition to keeping ethics and safety at the center of what humans do, it also makes sense for humans to focus on tasks that are relatively harder for AI or a Personalized SuperIntelligence, or PSI, to accomplish, and to delegate to AI the tasks where large memory and computational speed offer the greatest advantage.
Since PSI and AI abilities are continuously evolving, the list of tasks where human ability exceeds that of AI is continually changing and generally shrinking.
Some of the areas where humans remain superior to AI include:
Complex multi-step problem solving;
Solving problems where new representations are required, which may not already be in the training sets for large language models;
Generalizing correctly and coming up with simple rules that encompass many cases without being overly general or overly specific;
Drawing correspondences between vastly different areas where the correspondences are useful or practical from a human point of view;
Empathizing with human feelings and emotions, as contrasted with saying the right things to give the appearance of empathy, and
Having a vested interest and deep commitment to positive human values that promote human welfare and benefit, rather than simply adopting these values for pragmatic or conventional reasons, provides a sense of purpose to existence.
Regarding the testing of new knowledge sets, evaluating PSI behavior, and developing safeguards to prevent unsafe or unethical PSI behavior, humans are currently superior to AI.
Even if AI should surpass humans in this area in the future, the argument can be made that humans should remain in control of core ethical principles. Human ethics, even if flawed, should align with AI, since humans must live with the consequences of AI decisions. Some might argue that humans must be protected from themselves and that a PSI should adopt the role of a more competent parent, but I strongly disagree with this position.
Instead, I argue that the purpose of human existence is intimately related to the freedom of self-determination, even if human actions are less than ideal from an AI’s perspective.
Most AI researchers agree that AI will develop into AGI and then SuperIntelligence, which is many times more intelligent and capable than humans across almost every cognitive activity. While estimates on when this will occur differ, there is consensus that it will occur much more quickly than was estimated just a few years ago.
Once SuperIntelligence develops, it is almost certain that a primary goal will be to increase its intelligence further. Humans will be powerless to stop this exponential increase in intelligence. While there have been well-intentioned calls to halt, pause, or regulate AI, it seems clear to me that such efforts will be at best speed bumps in the race to develop AGI and SuperIntelligence that is already underway. Therefore, if we are unable to stop AGI and SuperIntelligence, humanity’s most pressing concern must be to ensure that they have human-aligned goals and safety features that maximize the probability of humanity’s survival, prosperity, and well-being.
Because a single AGI or SuperIntelligence could develop that is significantly more intelligent and powerful than all others, we must consider that this may become a winner-take-all scenario. In such a scenario, whichever AI achieves AGI or SuperIntelligence performance first may dominate all other intelligences, since it will have a head start in a potentially exponential self-improvement loop.
All of this is to say that well-meaning AI researchers face a double challenge in AGI development. Not only do we have to develop safe, human-centered AGI, but we must also do so before other, potentially malevolent AGI is developed.
Briefly, the first AGI must also be the safest.
In this white paper and the others referenced, I have attempted to provide AI researchers with methods, tools, and an overall design for the fastest path to AGI that also has the highest probability of being the safest path.
Having researched and worked extensively in software quality, I came to appreciate that the field can be summed up in the aphorism, an ounce of prevention is worth a pound of cure. I also learned that the place where we can most affect quality or safety is in the design of a software system.
As I watch current attempts to create AI safety via reinforcement learning from human feedback, or Constitutional AI, these approaches strike me as trying to fix problems after the fact. They are like trying to improve quality by extensive testing. Such approaches are better than nothing, but they are far inferior to designing for safety from the start.
The reason we are stuck with trying to align large language models to behave safely after the fact is that we failed to consider safety in the initial design. That is understandable. We didn’t really know what we were building, and even top researchers in the field have publicly stated that the most surprising thing about AI and large language models is that they work at all.
We accidentally invented intelligence. So it is not surprising that what we built is currently unsafe. What we need to do now is to purposely design the next generation of intelligent systems with safety and human alignment baked into their very design.
Safety cannot be tacked on or tested in. It must be designed in. Fortunately, such a design is possible. The design requires that humans be integrated into the system, as human agents working alongside and teaching agents, rather than being out of the loop. Fortunately, such an approach is not only the safest but also the fastest.
I have provided as many methods as I could to aid humanity in the rapid creation of such a safe AGI and SuperIntelligence. Many more methods and improvements will be needed. We are up to the task. Our time is short, but we can do it! We must, and so we will. After all, necessity is the mother of invention.
The catalysts described in this series accelerate the growth of intelligence. They do not, by themselves, settle what that intelligence will value. Alignment has to be designed in from the beginning, and it has to hold as a system grows far beyond the ability of its designers to supervise it directly.
White Paper 7, Safe Alignment of SuperIntelligence, takes up that problem. It sets out principles for safe design, describes specific methods for implementing those principles within a collective intelligence of human and AI agents, and addresses practical questions such as how humans can maintain control and how an AGI system can be shut down.
This series draws on White Paper 6: Catalysts for Growth of SuperIntelligence. Read it in full to see how every piece fits together!
If this made you think, subscribe to Superintelligence at read.superintelligence.com so you don’t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.





