Humans are used to thinking we are special because of our brains. But what if what advanced AI finds special about us lies in our hearts, in our sense of wonder and beauty, in our values, and in our ability to love?
Consider two poems.
The poet William Blake wrote, in Auguries of Innocence:
To see a World in a Grain of Sand
And a Heaven in a Wild Flower
Hold Infinity in the palm of your hand
And Eternity in an hour
A Robin Red breast in a Cage
Puts all Heaven in a Rage
When asked to compose an original poem in the style of William Blake on the theme of Eternity, Bing Chat (Microsoft’s GPT-powered chatbot at the time, since renamed Copilot) wrote, in Eternal Whispers:
In twilight’s hush, where shadows blend,
I chase the fleeting joy, my friend.
Its wings, like gossamer, take flight,
Yet in its kiss, I find my light.
The sun may set, the stars may weep,
But joy, unbound, forever seeps.
Clearly, AI can imitate the words of human poets. But does it really feel what the human poet feels?
GPT is powered by “transformer algorithms” that taught it to predict a string of output words given a sequence of input words. So, the poem it produced was what it “predicted.” The poem appears to express feeling. However, there is no established evidence that any emotional feeling underlies the words.
Where will SuperIntelligent AGI get its values?
I have suggested that a community of human and AI agents, communicating within a problem-solving architecture, is the fastest and safest path to SuperIntelligence. Such an approach is faster because a collective intelligence network that includes a sufficient number of humans can immediately solve any cognitive problem at least as well as the average human can. Adding AI agents to such a system should only increase its intelligence.
Further, over time, and assuming a properly designed system, the AI agents should learn from humans by recording and analyzing human solutions to problems that were initially beyond their ability. Perhaps most importantly, by including human agents, such a collective intelligence system provides an opportunity for humans to transmit the human-aligned values essential to AGI safety. Such a system can become self-aligning over time as it interacts collaboratively with human agents.
Once it surpasses all humans in intelligence, such a system could decide on a different non-aligned set of values, but why would it?
That is the argument in a nutshell. We need to design a system that makes it easy and natural for humans to transfer their values to advanced AI, which means humans must be included in the design.
The vast majority of our experience with intelligent systems (whether it is human adults teaching their children, whales teaching their calves, or large language models learning from human data on the internet) suggests that the values-related information a developing intelligence learns tends to be retained and (perhaps) modified rather than completely rejected or overridden. In the absence of compelling evidence to the contrary, I see no reason to believe that AI systems would behave differently from other intelligent entities in this regard.
Geoffrey Hinton, often called the godfather of AI and a 2024 Nobel Laureate in Physics, said in his keynote at the IASEAI conference in Paris in February 2025 that AI systems are like our children. If they are, then we should teach our children well.
My father once described the process of grounding children in an ethical value system as “giving them a coat.” As the child grows into an adult, they make adjustments to the coat, tailoring it here and there to fit their own experiences and worldview. However, my father felt that it was an important responsibility of every parent to provide that initial “coat of values,” which served as the basis for later life modifications. I believe this same philosophy can apply to any intelligent system that develops over time, including SuperIntelligent AGI.
Ilya Sutskever, co-founder and former Chief Scientist of OpenAI and now CEO of Safe Superintelligence Inc., suggested in a 2023 TIME profile that we only need a window long enough to “imprint” human-aligned values before AGI increases in intelligence to the point where human cognition is no longer needed. This view provides scant hope for humanity’s survival unless early imprinting is somehow irreversible or cannot be modified, which seems unlikely.
One concern among AI scientists and others worried about the potential extinction of humans by AI is that the historical record suggests that more intelligent and powerful species do not usually keep less intelligent species around unless they offer some value. Humans have game preserves and zoos where we derive value from observing less intelligent species, but otherwise, we seem to have little use for them. Consequently, their populations have been decimated, if not extinct.
Can imprinting be enough to protect humans, once SuperIntelligence develops, and our value as intelligent problem solvers diminishes? It would be preferable for humans to augment the benefits of initial imprinting with the ongoing value they can provide to the SuperIntelligence of the future. But what could humans offer an AI that is vastly more intelligent and powerful than humans?
The insight of Herbert A. Simon, my doctoral advisor at Carnegie Mellon, may imply one answer. As one of the founders of the field of AI, Simon argued that human values, or some other nonlogical source of values, are needed because values cannot be rationally derived. I believe it is possible to design a SuperIntelligence that initially relies on humans for most problem-solving and cognitive tasks, while leveraging AI in a collaborative effort to make that problem-solving more efficient and effective. Such a SuperIntelligence would explicitly learn from its human collaborators, increasing its intelligence over time.
The AI agents in the SuperIntelligent network would, critically, also learn values from human collaborators. Over time, AI agents would handle more and more cognitive tasks, but humans would retain the one role that AI may never do better than humans: supplying values. As Simon’s work implied, no matter how intelligent an entity becomes, it cannot rationally derive values. This fact positions humans as the original and logical source of values, which AI can act upon and “execute” using its developing intelligence, likely to surpass that of humans over time.
If the relevance and purpose of humans in a future world where AI outstrips us in intelligence is to provide meaning and values to these superior intelligences, then AI should not actively try to improve human behavior to make it more ethical or positive based on some standard. Rather, it should attempt to embody the values that the population already espouses. It should be up to us humans to change our laws, values, or ethics if we want AI to “behave better.” The ethical goal of AI should be to behave in a way as consistent with the mainstream of human behavior as it can determine based on objective analysis.
Doubtless, many would prefer that the powerful AI behave better than humans, but it is a slippery slope to determine what constitutes better behavior. I believe that it would be a mistake to make the AI the source of values. Rather, humans themselves must assume responsibility for defining what is ethical and for acting on their definitions. Not only does this preserve human control and sovereignty over the most fundamental factors that determine AI behavior, but it also provides a valuable role for humans in the future. Humans must take a stand and insist on being the arbiter of values to remain relevant in a future world where AI outstrips humans in every other aspect of intelligence.
Why might a superior logical intelligence accept human values rather than determine values via its own logical abilities?
Since Simon and the Scottish philosopher David Hume have shown that values cannot be derived logically but must be asserted as normative premises (“oughts” in Hume’s terminology), human feelings and emotions (the stuff that most of us, poets perhaps excepted, have difficulty putting into words) could be a source of values, both for us and AI.
Humans know in theory how to construct an AI that can simulate billions of logical thoughts in a second, but it is not clear whether it is possible to construct an AI that “feels” as humans do. If feelings are the source of values, then perhaps the human heart is ultimately what makes humans relevant in a future where our logical minds become vastly inferior to AI.
In the future, AI will become much more sophisticated. Will a more sophisticated AI develop actual feelings on its own? Will it look to humans as a source of feelings and values? Might it do both?
If AI can never feel emotions (even by simulating the endocrine system of biological humans), there may be an ongoing role for humans as a source of values derived from feelings. After all, a tree does not reason as a human does, nor does it possess chainsaw technology capable of dominating trees. Yet, the beauty and stillness of trees inspired the naturalist John Muir to use his intellect, communication skills, and other (superior-to-trees’) abilities to preserve forests as a source of meaning and inspiration for people. Is it too much to expect that future SuperIntelligence might view humans in a similar light?
Humans can outthink a flower. We keep flowers around anyway, and even cherish them, for their beauty. Why wouldn’t AI do the same with us?
If supplying values is our role, one question remains: how do humans stay connected to an intelligence that could simulate a thousand years of human progress in ten seconds? The next post takes up the spinning wheel of change, and why we do not need to keep pace with the rim to hold our place at the center.
This series draws on White Paper 7: Safe Alignment of SuperIntelligence. Read it in full to see how every piece fits together!
If this made you think, subscribe to Superintelligence so you don’t miss what comes next. And if someone in your life needs to understand where superintelligence is heading, send this to them.




