Although there have been many well-intentioned calls to halt, pause, slow, or regulate AI development, unfortunately, there is little evidence of anything other than a speedup in the race to AGI.
In 2024, Max Tegmark, the MIT physicist and president of the Future of Life Institute, said, “an AGI race is a suicide race.” That’s because if anyone’s SuperIntelligence escapes human control and becomes malevolent or misaligned with human values, all of us lose.
In fact, there is a serious risk that such a SuperIntelligence could cause the human race to go extinct. Therefore, whether we like it or not, all of humanity is engaged in a race to find the safest and most aligned SuperIntelligence that is possible.
Pause or no pause, we have to find a way to align the values of advanced AI and to design it to look to us for those values.
Pausing slows down the trainwreck.
Preventing it requires understanding the core problem, and that problem comes into focus when we look at what we are building.
AI systems are not subject to the same cognitive constraints as humans. Specifically, the phenomenon of “bounded rationality” does not apply to AI systems. Bounded rationality was central to the research that led to Herbert A. Simon, my doctoral advisor at Carnegie Mellon and one of the founders of the field of AI, receiving the 1978 Nobel Prize in Economics. Or, more accurately, the limits of bounded rationality for an entity that can process trillions of times more information, trillions of times faster than a human, are so remote that compared to a human, such a system has effectively unbounded rationality.
Similarly, the perceptual limits that apply to humans (our limited range of vision, hearing, smell, taste, and touch) do not apply to AI systems that can detect all wavelengths of electromagnetic radiation, all frequencies of sound waves, “odors” far beyond the range of human (or even animal) perception, and pressures that are undetectable to humans as well as pressures that would instantly crush a human. Beyond AI’s superior range of perception, there is also the matter of its superior scope.
A human can see only what is directly in front of them, at a specific resolution, and over a relatively small distance. An AI can theoretically perceive everything that happens on Earth, including in the deep oceans and the high stratosphere, all simultaneously, with incredibly precise resolution (think electron microscopes) and across extreme distances (think James Webb Space Telescope). Such capabilities of relatively “unbounded perception,” combined with “unbounded rationality,” enable SuperIntelligence far beyond what humans can easily comprehend, let alone emulate.
We label such potential entities with words and phrases like “SuperIntelligence,” “Artificial Super Intelligence,” or “Super Intelligent AGI.”
But such labels fail to capture the substantial difference in intelligence potential we are trying to explain.
Geoffrey Hinton, the computer scientist often called the godfather of AI and a 2024 Nobel Laureate in Physics, has compared humans to two-year-old children trying to outsmart an adult.
Others have suggested our limited human intelligence is like that of a pet compared to its human master.
I have suggested that the difference in intelligence may be analogous to that between an amoeba and Albert Einstein, with humans as the amoeba.
All of these analogies probably fall short of the eventual reality.
How can humans have any guarantee that such a vastly superior SuperIntelligence will have interests that are aligned with those of humans?
It’s a huge existential risk with an innocuous-sounding name: the alignment problem. Unfortunately, simply naming the problem does little to solve it. However, Simon had an idea more than forty years ago that might help us.
Simon wrote a relatively obscure book entitled Reason in Human Affairs (1983). In contrast to the nearly 1,000 pages he wrote with fellow AI pioneer Allen Newell on Human Problem Solving, Reason in Human Affairs is a mere 115 pages. Moreover, it is highly readable and easy for the average high school student to understand.
Yet within the pages of this remarkable little book, Simon reminds us of an essential idea that might hold the key to solving the alignment problem. It appears in just two sentences, at the bottom of page 7 of Simon’s little book:
We see that reason is wholly instrumental.
It cannot tell us where to go; at best it can tell us how to get there.
That’s it. Just twenty-four words!
But it means that there is no rational, logical way to derive what is right and what is wrong.
It’s a restatement of the argument, made in 1740 by the Scottish philosopher David Hume, that moral statements (“oughts”) cannot be derived from empirical facts (“is’s”). While some philosophers have debated the truth of this position, Simon agrees with it, stating that:
None of the rules of inference that have gained acceptance are capable of generating normative outputs purely from descriptive inputs. The corollary to ‘no conclusions without premises’ is ‘no oughts from is’s alone.’
How does that help us with the alignment problem?
If Simon and Hume are correct in their thinking, a SuperIntelligent AGI will be no better than humans at coming up with right and wrong. For all its superior processing speed and perception, SuperIntelligence will still run up against the hard fact that there is no way to derive morality, no matter how intelligent it becomes rationally. I suggest that this is a good thing for our species.
If we accept that values cannot be derived logically, then we are left with the question: Where will SuperIntelligent AGI get its values?
One source of these values could be the humans who initially created the SuperIntelligence. To increase the likelihood of this happening, AI researchers and engineers must design systems that maximize the transfer of human-centered values to SuperIntelligent AGI.
The super smart AI of the future will still have no logical way to determine right from wrong. It needs us humans to be the source of its values.
That is the core problem, and we need to understand it before calling for regulation or a pause. Pausing may slow the trainwreck. Preventing it requires designing advanced AI to look to humans for its values.
I have described the safest and fastest path to SuperIntelligence as arising from the collective intelligence of many agents, human and AI, working together. In contrast to the Mixture of Experts approach, which combines separate AI models trained in specific domains, that path includes human experts alongside AI experts.
The next post shows how such a design gives humans a natural way to pass their values to advanced AI, the way parents give their children an initial coat of values, and why the human heart may be what keeps us relevant to a SuperIntelligence that far outthinks us.
This series draws on White Paper 7: Safe Alignment of SuperIntelligence. Read it in full to see how every piece fits together!
If this made you think, subscribe to Superintelligence at read.superintelligence.com so you don’t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.





