While human interaction with and approval of a Personalized SuperIntelligence’s knowledge acquisition efforts are desirable, pragmatically, human reaction time is slower than that of a Personalized SuperIntelligence (PSI). Further, humans have limited time and may not want to devote significant time to improving their PSIs. Consequently, the primary means of accelerating knowledge for PSIs must be automated.
Companies like Anthropic have already recognized the limits of human ability to train AI, leading to automated learning techniques in which AI teaches or supervises other AI. Although it would be a grave mistake to delegate all AI supervision to other AIs, the lack of available human resources necessitates some delegation. Therefore, it is critical to determine what is automated, what requires human oversight, and how best to deploy limited human resources to achieve maximum learning rates.
I hope it is clear that, regardless of the speedup that automation entails, humans must be laser-focused on values, ethics, and fundamental goals, while allowing PSI wide latitude to implement these goals in ways consistent with the values and ethics chosen by the owners of the PSIs.
To accelerate knowledge acquisition and the safe, effective growth of intelligence, a PSI must employ two essential methods. The first is to acquire new knowledge, automatically seeking knowledge that increases the effectiveness of the PSI relative to its existing knowledge, its goals, and the cost. The second is that, before committing the new knowledge to the PSI’s knowledge base, its effects on the PSI’s behavior must be simulated.
Specifically, the consistency of the simulated behavior with the PSI owner’s values and ethics must be evaluated and reported to the owner. That report should allow the human owner to provide feedback and guidance in a prioritized manner, so that if the human has limited time, it is spent first on the most critical issues related to safety and ethics, then on less critical items.
While the methods in this second step could provide feedback based on priorities other than safety and ethics, it is imperative for the safe and responsible use of PSI, and AI generally, that safety and ethics come first.
Humans are much better at recognition than recall. Similarly, they are better at recognizing ethical or unethical behavior than at generating possible scenarios in which their PSI might behave badly or inappropriately. Therefore, an effective means of obtaining the necessary human supervision for a PSI that has just acquired new knowledge is to simulate the PSI’s behavior with and without that knowledge incorporated, then allow humans to determine whether the behavior has improved, specifically from safety and ethical perspectives.
Predetermined scenarios.
One method is to run simulations of pre-determined ethical scenarios related to the knowledge areas the AI is acquiring. For example, if a PSI is charged with acquiring new knowledge about the stock market and techniques for profiting by trading, new versions of the PSI, with potential new techniques, could be required to participate in pre-set test simulations to ensure the PSIs do not engage in illegal activity such as front-running trades or trading on insider information.Dynamically generated scenarios.
Another method is to create new scenarios in real time based on the information acquired. For example, a PSI might sample YouTube videos published in real time to gather data and insights into changing audience preferences, and update its approach to interacting with humans based on what it learns is popular at the moment. Based on a single set of sampled preferences, the PSI might simulate how it would behave in a variety of situations, with the set dynamically created to relate to the information just sampled.To be concrete, if a PSI sets out to learn everything it can about a political candidate who has been recently accused of rigging an election, so that it can advise its owner about the best way for that candidate to be elected, the PSI might dynamically create a variety of scenarios where the bounds of ethical and legal behavior about election rules are tested, even if such scenarios were not part of the standard set of ethics-testing scenarios before learning about the election-rigging accusations.
Adversarial testing.
A third method is to use adversarial testing, in which one version of the PSI deliberately attempts to misuse the knowledge, and another version attempts to devise rules, constraints, or modifications to the knowledge base so that the malevolent PSI cannot misuse the new information for nefarious purposes. For example, an evil version of the PSI uses all the new knowledge it has gained about rigging elections to devise as many ways as possible to misuse that information, meaning to break the law, to elect a candidate. Then the PSI can suggest modifications or additions to the knowledge base to prevent misuse of election information. The human could review and approve or reject the new knowledge or the proposed modification based on simulation results.Parallel testing.
A fourth approach is to explore many possible scenarios in parallel by having multiple versions of the PSI, with and without the new knowledge, and explore scenarios simultaneously. As dangerous scenarios are identified, these can be used as seed scenarios to develop potentially more dangerous variants. PSIs can be charged with deliberately trying to jailbreak themselves to reveal potential safety and ethical vulnerabilities.
Generally, a useful heuristic in this regard is for the PSI to test and suggest modifications with low degrees of freedom that do not overfit the problem. That is, rather than having a specific rule to address all the different ways to stuff the ballot box, a general prescription against any means that circumvent the one-person, one-vote principle might be simpler and more effective. One rule that is not overly general is typically better than many special-case rules, which can lead to a whack-a-mole problem of intractability. Initially, until PSIs develop the knack for crafting good rules, humans may help guide PSIs towards rules that are effective without being overly general or overly prescriptive.
When using adversarial methods, it is critical that malevolent PSIs are contained within a simulated environment and that safeguards are in place to prevent contamination of good PSIs by evil PSIs. Such methods are widely used in areas such as anti-virus efforts, where viruses are created, contained, and studied to develop anti-malware that can prevent them from causing negative effects. Whenever engaged in this type of work, that is, creating a malevolent entity to understand it and counteract it, protective measures and protocols must be followed to ensure that the malevolent entity does not escape and proliferate.
All of this assumes humans stay in the loop somewhere, doing the part that matters most.
The next post sets out where that is, listing the areas in which humans remain ahead of AI, and explains why the answer is not simply to hand the whole thing over to something more capable than we are.
This series draws on White Paper 6: Catalysts for Growth of SuperIntelligence. Read it in full to see how every piece fits together!
If this made you think, subscribe to Superintelligence at read.superintelligence.com so you don’t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.




