<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[SuperIntelligence]]></title><description><![CDATA[The fastest path to superintelligence is the safest path.]]></description><link>https://read.superintelligence.com</link><image><url>https://substackcdn.com/image/fetch/$s_!gyBu!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0dda15a0-44f6-46ec-92b3-dc2eaabed8df_256x256.png</url><title>SuperIntelligence</title><link>https://read.superintelligence.com</link></image><generator>Substack</generator><lastBuildDate>Fri, 07 Aug 2026 23:57:15 GMT</lastBuildDate><atom:link href="https://read.superintelligence.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Dr. Craig A. Kaplan]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[superintelligencebyiq@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[superintelligencebyiq@substack.com]]></itunes:email><itunes:name><![CDATA[Dr. Craig A. Kaplan]]></itunes:name></itunes:owner><itunes:author><![CDATA[Dr. Craig A. Kaplan]]></itunes:author><googleplay:owner><![CDATA[superintelligencebyiq@substack.com]]></googleplay:owner><googleplay:email><![CDATA[superintelligencebyiq@substack.com]]></googleplay:email><googleplay:author><![CDATA[Dr. Craig A. Kaplan]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The Four Phases of Safe AGI]]></title><description><![CDATA[A step-by-step design for training SuperIntelligence on the values of billions of people]]></description><link>https://read.superintelligence.com/p/the-four-phases-of-safe-agi</link><guid isPermaLink="false">https://read.superintelligence.com/p/the-four-phases-of-safe-agi</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Fri, 07 Aug 2026 12:49:16 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!NfMs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdeed27bd-5960-4b2a-863b-6601e5e6391f_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NfMs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdeed27bd-5960-4b2a-863b-6601e5e6391f_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NfMs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdeed27bd-5960-4b2a-863b-6601e5e6391f_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!NfMs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdeed27bd-5960-4b2a-863b-6601e5e6391f_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!NfMs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdeed27bd-5960-4b2a-863b-6601e5e6391f_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!NfMs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdeed27bd-5960-4b2a-863b-6601e5e6391f_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NfMs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdeed27bd-5960-4b2a-863b-6601e5e6391f_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/deed27bd-5960-4b2a-863b-6601e5e6391f_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1971454,&quot;alt&quot;:&quot;A diverse line of people, including a child and a woman in a wheelchair, passes glowing blue orbs up scaffolding levels around a giant unfinished amber lattice structure, where two workers place a bright blue core into a glowing ring at its center. Text: Built by Billions. Step by step, ordinary people can put humanity's values at the heart of SuperIntelligence. Superintelligence, Safe Scalable AGI Series, by Dr. Craig A. Kaplan.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/208640281?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdeed27bd-5960-4b2a-863b-6601e5e6391f_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A diverse line of people, including a child and a woman in a wheelchair, passes glowing blue orbs up scaffolding levels around a giant unfinished amber lattice structure, where two workers place a bright blue core into a glowing ring at its center. Text: Built by Billions. Step by step, ordinary people can put humanity's values at the heart of SuperIntelligence. Superintelligence, Safe Scalable AGI Series, by Dr. Craig A. Kaplan." title="A diverse line of people, including a child and a woman in a wheelchair, passes glowing blue orbs up scaffolding levels around a giant unfinished amber lattice structure, where two workers place a bright blue core into a glowing ring at its center. Text: Built by Billions. Step by step, ordinary people can put humanity's values at the heart of SuperIntelligence. Superintelligence, Safe Scalable AGI Series, by Dr. Craig A. Kaplan." srcset="https://substackcdn.com/image/fetch/$s_!NfMs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdeed27bd-5960-4b2a-863b-6601e5e6391f_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!NfMs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdeed27bd-5960-4b2a-863b-6601e5e6391f_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!NfMs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdeed27bd-5960-4b2a-863b-6601e5e6391f_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!NfMs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdeed27bd-5960-4b2a-863b-6601e5e6391f_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4><span>Consider a scenario that could be implemented by a company such as Meta. Meta has a tremendously valuable asset for implementing scalable, safe AI: its huge user base. The data contained and content posted by billions of users is certainly large enough to provide a representative, statistically valid sample of human values and ethics. Where certain geographies or populations are under-represented, a company of Meta&#8217;s size can form partnerships to access representative data for those populations. Everything this series has described, the community of teachers, the weighting schemes, the coverage of the most important scenarios, the spinning wheel with values at its center, can be assembled from assets a company like this already owns.</span></h4><ul><li><p><span>The starting point is a personalized AI agent for every user. Because Meta users have, in aggregate, posted enormous amounts of preference and content information, it is relatively easy to create such agents. </span></p></li><li><p><span>With the press of a button, a user can specify that a base LLM be trained on their personal data so it behaves more like them, reflecting their preferences, knowledge, skills, and ethical values. </span></p></li><li><p><span>The trained agent can immediately begin filtering unwanted content and suggesting postings on the user&#8217;s behalf. </span></p></li><li><p><span>By observing the user&#8217;s actions, it can fine-tune itself and learn more about its owner. </span></p></li><li><p><span>To the degree that users participate in immersive virtual environments, even richer data becomes available, since AI can observe every motion and behavior in the virtual world, a gold mine of information far richer than the standard internet datasets widely used today. </span></p></li><li><p><span>Ethical customization of this kind is primarily a function of large representative datasets, sufficient computing power, and powerful machine learning algorithms, all of which Meta possesses.</span></p></li></ul><h4><span>Customization does not need to start from zero. </span></h4><p><span>Knowledge modules, described in earlier white papers in this series, are essentially sets of training weights that can be combined with an LLM&#8217;s existing weights to change its behavior in known and predictable ways. Suppose the human owner of an AI is a devout Christian. One module might be a King James Bible package: an off-the-shelf LLM extensively trained on the Old and New Testaments, with the exact training corpus available and benchmarks describing how its behavior differs from the base model. The owner could start with the Bible-trained model and then further customize it, perhaps emphasizing the golden rule and New Testament ideas of charity, forgiveness, and mercy while minimizing passages that do not align with the owner&#8217;s modern ethical sensibilities. Customizing a pre-trained AI in this way saves the time and effort of starting from scratch.</span></p><h4><span>Modules also allow delegation. </span></h4><p><span>Groups of humans can delegate their ethical influence to a pre-trained module that represents their position well, so that the Bible-trained AI, for example, could vote on the ethical preferences of many humans at once. Delegation is inferior to each human explicitly training their own AI. Still, it may greatly increase the total number of humans represented in an AI&#8217;s value system, even if many are represented by proxy. In the preferred implementation, there is a marketplace for such modules, where users can share, trade, or license the packages they create, with users owning the data they generate and the platform serving as a broker. The result is maximum choice about which starting point best fits each person.</span></p><blockquote><p><strong><span>Building the base model that users then customize might follow four general phases: </span></strong></p><ol><li><p><strong><span>Training a base model,</span></strong></p></li><li><p><strong><span>Customizing it for each user,</span></strong></p></li><li><p><strong><span>Combining knowledge from many customized agents, and </span></strong></p></li><li><p><strong><span>Refining the agents through collective problem-solving. </span></strong></p></li></ol></blockquote><h4><span>In the first phase, the base LLM is trained on data reflecting the composition of its target users, then made safer through a large corpus of ethical and safety scenarios. </span></h4><ul><li><p><span>A variant of Constitutional AI, using a trusted earlier model under human oversight, efficiently covers the most common cases. </span></p></li><li><p><span>At the same time, users are offered incentives to crowdsource new safety scenarios, which trusted users and employees then filter to achieve coverage of the most frequent and impactful cases. </span></p></li><li><p><span>Redundant human feedback ensures that the ethics taught to the base model never rely on a single person&#8217;s input, and the model is released first to a select group whose feedback further refines its safety.</span></p></li></ul><h4><span>In the second phase, each user&#8217;s agent learns its owner&#8217;s ethics. </span></h4><ul><li><p><span>The agent draws on a ranked corpus of ethical questions, asking the most important ones for which data is still lacking before any others, since humans are generally unwilling to answer questions posed by AI. </span></p></li><li><p><span>The highest-ranked questions are the fundamental ones, such as under what circumstances, if any, it is appropriate for an AI agent to harm a human directly, or to disobey a direct order. </span></p></li><li><p><span>With the user&#8217;s permission, the agent can also learn passively, analyzing the user&#8217;s posted content with a single button press and authorizing tuning, with information weighted by recency, type, and source, as earlier posts described. </span></p></li><li><p><span>Users or the platform can also specify alert conditions, such as a major court decision or news event with ethical implications, as triggers for event-based updates to the training. </span></p></li><li><p><span>A user can even select existing customized agents as training sources, saying, in effect, &#8220;I want my AI trained using the weights of the agent belonging to the pastor of my church,&#8221; or choosing to have a Tibetan Monk AI do 80% of the tuning and a Humanistic Philosopher AI the remaining 20%. </span></p></li><li><p><span>Consistent with the principle that humans in the loop are the primary means of catching AI errors, input from humans is weighted more strongly than input from other AIs, and input from the owner is more strongly weighted than input from anyone else.</span></p></li></ul><h4><span>In the third phase, the customized agents combine what they have learned. </span></h4><ul><li><p><span>Weights from multiple agents can be merged on a one-vote-per-AI, one-AI-per-user basis to obtain a representative, statistically valid set of AI ethics that generalizes across all the humans represented. </span></p></li><li><p><span>In the broadest implementation, combining weights from as many agents as possible creates a value system that broadly represents the values of all humans on Earth, and such a value system would likely reduce the risk of AI-driven human extinction. </span></p></li><li><p><span>Alternatively, many customized agents can be presented with ethical dilemmas and vote on the best actions, optionally checking with their human owners before casting votes on important matters, with safeguards that trigger human review whenever conclusions conflict with generally accepted precepts such as valuing human life. </span></p></li><li><p><span>As AIs become increasingly intelligent, the burden of representing their owners will fall increasingly on the agents themselves, since humans cannot keep up with the speed of AI thought. </span></p></li><li><p><span>The guiding principle returns to the spinning wheel: human input is preserved on matters closest to the center, where alignment and purpose live. </span></p></li><li><p><span>As long as core human values, such as the golden rule and love for all humans, are preserved near the center, AI can make many decisions relatively autonomously on the wheel's periphery.</span></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vhVl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15c3c09b-796e-4d56-8882-ef34c167cdbb_1535x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vhVl!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15c3c09b-796e-4d56-8882-ef34c167cdbb_1535x1024.png 424w, https://substackcdn.com/image/fetch/$s_!vhVl!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15c3c09b-796e-4d56-8882-ef34c167cdbb_1535x1024.png 848w, https://substackcdn.com/image/fetch/$s_!vhVl!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15c3c09b-796e-4d56-8882-ef34c167cdbb_1535x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!vhVl!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15c3c09b-796e-4d56-8882-ef34c167cdbb_1535x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vhVl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15c3c09b-796e-4d56-8882-ef34c167cdbb_1535x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/15c3c09b-796e-4d56-8882-ef34c167cdbb_1535x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1382572,&quot;alt&quot;:&quot;all three arrows must terminate at the large circle, none may point at another agent, and the three arrows must not touch each other&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/208640281?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15c3c09b-796e-4d56-8882-ef34c167cdbb_1535x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="all three arrows must terminate at the large circle, none may point at another agent, and the three arrows must not touch each other" title="all three arrows must terminate at the large circle, none may point at another agent, and the three arrows must not touch each other" srcset="https://substackcdn.com/image/fetch/$s_!vhVl!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15c3c09b-796e-4d56-8882-ef34c167cdbb_1535x1024.png 424w, https://substackcdn.com/image/fetch/$s_!vhVl!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15c3c09b-796e-4d56-8882-ef34c167cdbb_1535x1024.png 848w, https://substackcdn.com/image/fetch/$s_!vhVl!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15c3c09b-796e-4d56-8882-ef34c167cdbb_1535x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!vhVl!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F15c3c09b-796e-4d56-8882-ef34c167cdbb_1535x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Each person trains one agent, and each agent counts once, so the merged values represent everyone.</figcaption></figure></div><h4><span>In the fourth phase, the agents, together with human agents, collaborate to solve real problems, following the principle that many heads are better than one. </span></h4><ul><li><p><span>Teams of agents can be assembled the way effective human teams are, by profiling skills against the task, seeking redundant coverage where the stakes are high, using reputation and track records, and letting market mechanisms match agents to work. </span></p></li><li><p><span>The more agents involved, the broader the collective range of knowledge and values, an advantage that matters because, as </span><a href="https://www.investopedia.com/terms/h/herbert-a-simon.asp"><span>Nobel Laureate Herbert A. Simon showed that information-processing constraints are a primary limit on any intelligence, human or artificial;</span></a><span> two safeguards run through everything. </span></p></li><li><p><span>The values being stored can be recorded on blockchain or other auditable structures, enabling verification that they are the values humans intended. </span></p></li><li><p><span>And because all problem-solving involves setting goals and subgoals, a series of ethics checks must be passed each time a new goal or subgoal is set, automating the enforcement of safety every time any agent on the network solves a problem, with those checks drawing on the representative values the community built and updating automatically as the community&#8217;s ethics evolve.</span></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!3jw8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dcad75-2438-4923-8b9f-a6a6ab205049_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!3jw8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dcad75-2438-4923-8b9f-a6a6ab205049_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!3jw8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dcad75-2438-4923-8b9f-a6a6ab205049_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!3jw8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dcad75-2438-4923-8b9f-a6a6ab205049_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!3jw8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dcad75-2438-4923-8b9f-a6a6ab205049_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!3jw8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dcad75-2438-4923-8b9f-a6a6ab205049_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/64dcad75-2438-4923-8b9f-a6a6ab205049_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1462459,&quot;alt&quot;:&quot;Diagram of four numbered phases: train a safe base model, customize one agent per person, combine what the agents learned, and solve problems together. Caption, if running one: The four phases run in sequence, from one safe base model to a community of humans and agents solving problems together.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/208640281?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dcad75-2438-4923-8b9f-a6a6ab205049_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Diagram of four numbered phases: train a safe base model, customize one agent per person, combine what the agents learned, and solve problems together. Caption, if running one: The four phases run in sequence, from one safe base model to a community of humans and agents solving problems together." title="Diagram of four numbered phases: train a safe base model, customize one agent per person, combine what the agents learned, and solve problems together. Caption, if running one: The four phases run in sequence, from one safe base model to a community of humans and agents solving problems together." srcset="https://substackcdn.com/image/fetch/$s_!3jw8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dcad75-2438-4923-8b9f-a6a6ab205049_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!3jw8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dcad75-2438-4923-8b9f-a6a6ab205049_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!3jw8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dcad75-2438-4923-8b9f-a6a6ab205049_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!3jw8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F64dcad75-2438-4923-8b9f-a6a6ab205049_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4><span>The implementation details serve a deeper purpose. </span></h4><p><span>Values, ethics, and domain knowledge are all constantly changing. Because values change most slowly, sitting closest to the center of the spinning wheel, they matter most for determining the behavior of SuperIntelligence. Even in a world where AI is trillions of times smarter than us, humans can retain a role as the source of values at the center of the rapidly evolving intelligence that is emerging. If humans can center our attention, thoughts, words, and actions on love, SuperIntelligent AI will perceive love, learn to love, and use its intelligence in the service of love. As the psychologist Viktor Frankl pointed out, humans have a driving need for purpose and meaning. We must design AI, AGI, and SuperIntelligent systems to have this need as well, and to look to humans to supply purpose and meaning. Our survival may depend upon it.</span></p><p><span>This series began with the failure of today&#8217;s safety training and ends with an integrated design: many human teachers, carefully weighted voices, coverage of what matters most, real-time safeguards, and human values held motionless at the center of an accelerating wheel. The design depends on assembling safe, capable AI agents and combining their judgment at scale, which makes the individual agents themselves the next question. As those agents grow into Personalized SuperIntelligences that far exceed their creators in intelligence, each becomes powerful enough to pose a serious threat to human safety, and testing alone cannot guarantee their safety. Safety must instead be built in by design. </span></p><p><strong><span>White Paper 5, </span><a href="https://www.superintelligence.com/whitepaper-5-personalized-si"><span>Safe Personalized SuperIntelligence</span></a><span>, takes up that problem, and so will the next series.</span></strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/the-four-phases-of-safe-agi/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/the-four-phases-of-safe-agi/comments"><span>Leave a comment</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/the-four-phases-of-safe-agi?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/the-four-phases-of-safe-agi?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div class="directMessage button" data-attrs="{&quot;userId&quot;:163471793,&quot;userName&quot;:&quot;Dr. Craig A. Kaplan&quot;,&quot;canDm&quot;:null,&quot;dmUpgradeOptions&quot;:null,&quot;isEditorNode&quot;:true}" data-component-name="DirectMessageToDOM"></div><div><hr></div><p><em><span>This series draws on </span><a href="https://www.superintelligence.com/whitepaper-4-scalable-agi"><span>White Paper 4: Safe, Scalable Artificial General Intelligence</span></a><span>. Read it in full to see how every piece fits together!</span></em></p><p><strong><span>If this made you think, subscribe to Superintelligence at </span><a href="http://read.superintelligence.com"><span>read.superintelligence.com</span></a><span> so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</span></strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/si-research-whitepapers&quot;,&quot;text&quot;:&quot;SuperIntelligence Design White Papers&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/si-research-whitepapers"><span>SuperIntelligence Design White Papers</span></a></p>]]></content:encoded></item><item><title><![CDATA[The Spinning Knowledge Wheel]]></title><description><![CDATA[AI must update fast-changing knowledge at the rim while human values hold still at the center.]]></description><link>https://read.superintelligence.com/p/the-spinning-knowledge-wheel</link><guid isPermaLink="false">https://read.superintelligence.com/p/the-spinning-knowledge-wheel</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Wed, 05 Aug 2026 13:03:53 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!X030!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c5f9c68-024a-4c21-851e-ce9566aefda3_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!X030!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c5f9c68-024a-4c21-851e-ce9566aefda3_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!X030!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c5f9c68-024a-4c21-851e-ce9566aefda3_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!X030!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c5f9c68-024a-4c21-851e-ce9566aefda3_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!X030!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c5f9c68-024a-4c21-851e-ce9566aefda3_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!X030!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c5f9c68-024a-4c21-851e-ce9566aefda3_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!X030!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c5f9c68-024a-4c21-851e-ce9566aefda3_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9c5f9c68-024a-4c21-851e-ce9566aefda3_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2237836,&quot;alt&quot;:&quot;A glowing gyroscope-like wheel spins against a dark background, its amber rim blurred with speed while a single bright blue core at the center holds perfectly still. Text: What Must Never Move. AI should update everything except the human values holding the wheel together. Superintelligence, Safe Scalable AGI Series, by Dr. Craig A. Kaplan&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/208397071?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c5f9c68-024a-4c21-851e-ce9566aefda3_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A glowing gyroscope-like wheel spins against a dark background, its amber rim blurred with speed while a single bright blue core at the center holds perfectly still. Text: What Must Never Move. AI should update everything except the human values holding the wheel together. Superintelligence, Safe Scalable AGI Series, by Dr. Craig A. Kaplan" title="A glowing gyroscope-like wheel spins against a dark background, its amber rim blurred with speed while a single bright blue core at the center holds perfectly still. Text: What Must Never Move. AI should update everything except the human values holding the wheel together. Superintelligence, Safe Scalable AGI Series, by Dr. Craig A. Kaplan" srcset="https://substackcdn.com/image/fetch/$s_!X030!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c5f9c68-024a-4c21-851e-ce9566aefda3_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!X030!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c5f9c68-024a-4c21-851e-ce9566aefda3_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!X030!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c5f9c68-024a-4c21-851e-ce9566aefda3_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!X030!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9c5f9c68-024a-4c21-851e-ce9566aefda3_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4><span>All knowledge is a moving target, with more recent knowledge generally superior to and supplanting earlier knowledge. At one time, the generally accepted view was that the world was flat. Today, almost everyone agrees that Earth looks much more like a sphere, and an AI trained on the flat-Earth view would be outdated and would need to update its knowledge.</span></h4><p><span>While scientific views, such as the shape of the Earth, may change very slowly, other forms of knowledge may change much more frequently. This is especially true of subjective ethical norms, where views the mainstream held confidently within living memory are now considered wrong. Norms of that kind can shift within a generation or less, while other forms of knowledge may remain valid for centuries. An efficient AI, therefore, needs a mechanism for updating its knowledge at the appropriate frequency, based on the velocity of change of the information, and one important dimension along which AI can categorize knowledge is the rate at which human opinions about the topic have changed.</span></p><h4><span>One way to think about this is to consider the difference between ethical principles and the fashions or interpretations of those principles. </span></h4><p><span>Humans have a long-standing principle that human life is valuable and should not be taken lightly. This general principle has survived for many thousands of years. It is incorporated into the laws and religious and moral traditions of almost all human groups, even though exceptions are made for war and certain other circumstances. On the other hand, certain ethical norms are more akin to fashions, which change depending on the group of humans being asked or the time at which they are asked. Affirmative action in college admissions served as an ethical norm for decades, until a Supreme Court decision began to influence it, and almost immediately, many companies and other organizations adjusted their norms, decision-making, and communication practices to align with the new mainstream view. Many other social attitudes are similarly fluid, changing far more quickly than long-lasting and widely accepted ethical precepts such as &#8220;thou shalt not kill.&#8221;</span></p><h4><span>The rate of change tells AI how to weight what it learns. </span></h4><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cagO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46417a28-3db6-487c-a197-3b88bbffd467_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cagO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46417a28-3db6-487c-a197-3b88bbffd467_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!cagO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46417a28-3db6-487c-a197-3b88bbffd467_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!cagO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46417a28-3db6-487c-a197-3b88bbffd467_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!cagO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46417a28-3db6-487c-a197-3b88bbffd467_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cagO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46417a28-3db6-487c-a197-3b88bbffd467_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/46417a28-3db6-487c-a197-3b88bbffd467_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1416081,&quot;alt&quot;:&quot;Chart showing how much weight AI gives knowledge over time: a flat blue line for fundamental values, a gently declining white line for slow-changing knowledge, and a steeply falling amber line for fast-changing knowledge. Headline: Not all knowledge ages at the same speed. Bottom line: The faster knowledge changes, the faster old knowledge loses its weight.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/208397071?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46417a28-3db6-487c-a197-3b88bbffd467_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Chart showing how much weight AI gives knowledge over time: a flat blue line for fundamental values, a gently declining white line for slow-changing knowledge, and a steeply falling amber line for fast-changing knowledge. Headline: Not all knowledge ages at the same speed. Bottom line: The faster knowledge changes, the faster old knowledge loses its weight." title="Chart showing how much weight AI gives knowledge over time: a flat blue line for fundamental values, a gently declining white line for slow-changing knowledge, and a steeply falling amber line for fast-changing knowledge. Headline: Not all knowledge ages at the same speed. Bottom line: The faster knowledge changes, the faster old knowledge loses its weight." srcset="https://substackcdn.com/image/fetch/$s_!cagO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46417a28-3db6-487c-a197-3b88bbffd467_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!cagO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46417a28-3db6-487c-a197-3b88bbffd467_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!cagO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46417a28-3db6-487c-a197-3b88bbffd467_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!cagO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F46417a28-3db6-487c-a197-3b88bbffd467_1536x1024.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Fundamental values keep their weight while fast-changing knowledge is steadily replaced by newer information.</figcaption></figure></div><p><span>The more rapidly a knowledge area changes, the more frequently updates should be made, and the more weight recent information should receive relative to older information. In areas where change is rapid or exponential, recent knowledge deserves proportionally greater weight. In areas where knowledge changes slowly and steadily, the advantage of recency is smaller. And where knowledge has been constant for long periods, as with firmly established principles like the high value of human life, new information should carry similar weight to old, since the principle itself is not moving. Underneath all of these adjustments lies a general rule: a core role of any intelligent system is to maintain an accurate representation of the current state of the world, and more recent knowledge generally describes that state better than older knowledge.</span></p><h4><span>The frequency of updates will only grow more important. </span></h4><p><span>With AI accelerating scientific discovery and technological change, it is conceivable that knowledge about the world will change faster than humans can update their collective understanding. Given the limitations of human information processing and our tendency to cling to outdated paradigms, knowledge may already be increasing far faster than most humans can comprehend or adapt. That said, when it comes to fundamental ethical principles like the value of human life, we are fortunate that these change relatively slowly. The interpretation and application of ethical principles may change with technological developments, but the principles themselves remain relatively constant.</span></p><blockquote><p><strong><span>One might imagine a spinning wheel with fundamental human values, such as love and the value of human life, near its center. At the very center of the wheel, the ethical principles are constant and motionless, just as the center of a spinning wheel does not move at all. The farther along the spokes one travels toward the rim, the faster the rate of change. All knowledge, including ethical knowledge, can be characterized as lying closer to the center or farther out on the rim. An efficient AI needs to update the areas on the rim very frequently, without changing the core human values to which those areas are relevant.</span></strong></p></blockquote><p><span>Humans cannot keep pace with the change at the rim of the spinning knowledge wheel. However, we can understand and orient the entire wheel by serving as its relatively slower-moving center, where the values and purpose of AI reside. Human-centered aligned AI must put relatively constant and fundamental human values at the center, while updating the knowledge closer to the rim faster than humans can conceive. This structure ensures that human values remain the center of AI systems that may become potentially trillions of times more powerful and knowledgeable than any one human, and it is essential if humans are not only to survive but also to prosper in the age of such systems.</span></p><p><span>The wheel also revisits the weighting question from earlier in this series. Long-lasting, fundamental knowledge should carry greater weight and be more resistant to change than short-term, fashionable opinions. One way to determine how fundamental an ethical precept is would be to actively survey people and ask them to rate it relative to other candidate precepts, while another is to passively analyze records of human behavior and draw conclusions from them. Passive analysis is more efficient, and active engagement is necessary to ensure the conclusions drawn from it are correct from a human perspective, so both methods are likely to be useful.</span></p><h4><span>Not all knowledge is a matter of opinion, however.</span></h4><p><span>Values and ethics are more the exception than the rule in this regard, since aside from artistic judgments, political and religious views, and other subjective areas, most human knowledge is factual. AI will likely want to weight knowledge that is factually accurate and justified by converging evidence more highly than unsubstantiated opinions on factual matters. While some people still believe the Earth is flat, this view should not be given equal weight to the spherical view, which is supported by a vast body of converging scientific evidence. The problem is tricky because humans tend to select facts that support their views, and the facts themselves change. At one time, not so long ago, the consensus medical opinion was that cigarette smoking was healthy for the lungs. To navigate these issues, AI must rely primarily on the scientific method, seeking valid, reproducible evidence and converging results, and applying tested tools such as Occam&#8217;s Razor, before accepting facts.</span></p><h4><span>However, AI must not confuse facts with values, or as the philosopher </span><a href="https://en.wikipedia.org/wiki/Is%E2%80%93ought_problem"><span>David Hume</span></a><span> put it, &#8220;is&#8221; with &#8220;ought.&#8221; </span></h4><ul><li><p><span>Values are necessarily subjective. </span></p></li><li><p><span>Arguments claiming that values are objective, such as the claim that everyone would agree that something causing all humans the most extreme misery imaginable is bad, are naive and fail to grasp that other, non-human entities might not accept such values as self-evident at all.</span></p></li><li><p><span>AI operates at the most fundamental level in a precise, logical manner. </span></p></li><li><p><span>It is all zeros and ones at the machine level.</span></p></li><li><p><span>To expect such a system to intuit somehow that human values are fundamental, or, worse, to expect it to derive human-centered values logically, is the worst kind of sloppy thinking. This kind can lead to human extinction. </span></p></li><li><p><span>Both David Hume and Nobel Laureate Herbert A. Simon had it right when they emphasized that there is no rational way to derive values. Rationality cannot tell us where to go; at best, it can tell us how to get there. </span></p></li><li><p><span>To delegate the destination, the fundamental subjective values that AI adopts, to AI itself, expecting it to determine right and wrong rationally, is sheer folly and must be avoided at all costs!</span></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vEe_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f48c26-e2c6-47e6-b2d6-00bafd96ab06_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vEe_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f48c26-e2c6-47e6-b2d6-00bafd96ab06_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!vEe_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f48c26-e2c6-47e6-b2d6-00bafd96ab06_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!vEe_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f48c26-e2c6-47e6-b2d6-00bafd96ab06_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!vEe_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f48c26-e2c6-47e6-b2d6-00bafd96ab06_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vEe_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f48c26-e2c6-47e6-b2d6-00bafd96ab06_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/44f48c26-e2c6-47e6-b2d6-00bafd96ab06_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2006279,&quot;alt&quot;:&quot;Illustration of a glowing driverless car with a powerful engine and an empty driver's seat, facing five unmarked diverging roads, while a human carrying a glowing map walks toward the open door. Labels read: AI can drive. Humans choose where. Headline: Rationality is the engine, values are the destination. Bottom line: The smartest engine cannot pick the destination.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/208397071?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f48c26-e2c6-47e6-b2d6-00bafd96ab06_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Illustration of a glowing driverless car with a powerful engine and an empty driver's seat, facing five unmarked diverging roads, while a human carrying a glowing map walks toward the open door. Labels read: AI can drive. Humans choose where. Headline: Rationality is the engine, values are the destination. Bottom line: The smartest engine cannot pick the destination." title="Illustration of a glowing driverless car with a powerful engine and an empty driver's seat, facing five unmarked diverging roads, while a human carrying a glowing map walks toward the open door. Labels read: AI can drive. Humans choose where. Headline: Rationality is the engine, values are the destination. Bottom line: The smartest engine cannot pick the destination." srcset="https://substackcdn.com/image/fetch/$s_!vEe_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f48c26-e2c6-47e6-b2d6-00bafd96ab06_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!vEe_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f48c26-e2c6-47e6-b2d6-00bafd96ab06_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!vEe_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f48c26-e2c6-47e6-b2d6-00bafd96ab06_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!vEe_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F44f48c26-e2c6-47e6-b2d6-00bafd96ab06_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Rationality can take AI anywhere, and only human values can say where to go.</figcaption></figure></div><p>An earlier post in this series promised to return to the question of why greater intelligence does not give AI the authority to determine humanity&#8217;s values. </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;77f15f62-1843-40b4-9c35-0b589e392e09&quot;,&quot;caption&quot;:&quot;Imagine millions of individuals, each customizing a personal AI agent and teaching it their knowledge and their values.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;sm&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Who Should Have the Most Influence on AI?&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:163471793,&quot;name&quot;:&quot;Dr. Craig A. Kaplan&quot;,&quot;bio&quot;:&quot;Dr. Craig A. Kaplan is CEO of iQ Company and founder of SuperIntelligence.com, focused on AGI and SuperIntelligence. Designs systems using collective intelligence and quantitative modeling. PhD, CMU; co-authored with Herbert A. Simon.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/039428e7-3d69-49db-b948-64d9fa4e36d9_256x256.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-07-24T12:46:18.970Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!COYW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3242eda0-262c-4ca1-9e0d-bf62e845c8ac_1536x1024.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://read.superintelligence.com/p/who-should-have-the-most-influence&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:207873900,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:8390910,&quot;publication_name&quot;:&quot;SuperIntelligence&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!gyBu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0dda15a0-44f6-46ec-92b3-dc2eaabed8df_256x256.png&quot;,&quot;belowTheFold&quot;:true,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p><span>The wheel is the answer. Intelligence operates along the spokes and out at the racing rim, while the destination sits at the motionless center, which belongs to humans. The final post in this series shows how a real company could put this entire design into practice, phase by phase, with humans supplying the values and the purpose at every step.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/the-spinning-knowledge-wheel/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/the-spinning-knowledge-wheel/comments"><span>Leave a comment</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/the-spinning-knowledge-wheel?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/the-spinning-knowledge-wheel?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><p><span>This series draws on </span><em><a href="https://www.superintelligence.com/whitepaper-4-scalable-agi"><span>White Paper 4: Safe, Scalable Artificial General Intelligence</span></a><span>. </span></em><span>Read it in full to see how every piece fits together!</span></p><p><strong><span>If this made you think, subscribe to Superintelligence at </span><a href="http://read.superintelligence.com"><span>read.superintelligence.com</span></a><span> so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</span></strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/si-research-whitepapers&quot;,&quot;text&quot;:&quot;SUPERINTELLIGENCE DESIGN WHITE PAPERS&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/si-research-whitepapers"><span>SUPERINTELLIGENCE DESIGN WHITE PAPERS</span></a></p>]]></content:encoded></item><item><title><![CDATA[How AI Can Catch Danger in Real Time]]></title><description><![CDATA[When AI is unsure, it stops and asks.]]></description><link>https://read.superintelligence.com/p/how-ai-can-catch-danger-in-real-time</link><guid isPermaLink="false">https://read.superintelligence.com/p/how-ai-can-catch-danger-in-real-time</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Mon, 03 Aug 2026 12:59:25 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!5IYy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99381244-a413-4a0e-8e62-f60d1c81bf91_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5IYy!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99381244-a413-4a0e-8e62-f60d1c81bf91_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5IYy!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99381244-a413-4a0e-8e62-f60d1c81bf91_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!5IYy!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99381244-a413-4a0e-8e62-f60d1c81bf91_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!5IYy!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99381244-a413-4a0e-8e62-f60d1c81bf91_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!5IYy!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99381244-a413-4a0e-8e62-f60d1c81bf91_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5IYy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99381244-a413-4a0e-8e62-f60d1c81bf91_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/99381244-a413-4a0e-8e62-f60d1c81bf91_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1128172,&quot;alt&quot;:&quot;A blue icon stream diverts one item into a glowing amber pause symbol surrounded by human reviewers, then returns it to the flow. Text reads &#8220;The Split-Second Safety Review&#8221; and &#8220;The safest answer is sometimes a pause.&#8221;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/208291735?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99381244-a413-4a0e-8e62-f60d1c81bf91_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A blue icon stream diverts one item into a glowing amber pause symbol surrounded by human reviewers, then returns it to the flow. Text reads &#8220;The Split-Second Safety Review&#8221; and &#8220;The safest answer is sometimes a pause.&#8221;" title="A blue icon stream diverts one item into a glowing amber pause symbol surrounded by human reviewers, then returns it to the flow. Text reads &#8220;The Split-Second Safety Review&#8221; and &#8220;The safest answer is sometimes a pause.&#8221;" srcset="https://substackcdn.com/image/fetch/$s_!5IYy!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99381244-a413-4a0e-8e62-f60d1c81bf91_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!5IYy!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99381244-a413-4a0e-8e62-f60d1c81bf91_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!5IYy!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99381244-a413-4a0e-8e62-f60d1c81bf91_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!5IYy!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F99381244-a413-4a0e-8e62-f60d1c81bf91_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>AI does not need to anticipate every danger in advance to stay safe. </h4><h4>Human and AI agents can dynamically flag potential ethical issues in real time as they are encountered, then present them to other groups of agents for resolution. Rather than relying on experts or crowdsourcing to determine the full range of ethical scenarios in advance, the real-time flagging approach allows AI, AGI, and SuperIntelligent systems to detect potential issues and pause work until additional human input helps the system determine the ethical approach.</h4><p>Of course, in time-critical situations, pausing or delaying might not always be possible, but the approach can be used for many issues that do not demand an immediate response. Including this dynamic approach of delaying responses until ethical input is received can reduce an otherwise exponential space of possibilities to a manageable size. One implication is that critical, high-stakes issues requiring an immediate response, such as whether to launch a counterattack to a perceived missile launch or other military applications, will require proportionally more path coverage and training in advance than situations where a delay in response is acceptable.</p><h4>Constitutional AI approaches are generally suboptimal for establishing ethical knowledge bases, partly because they rely on rules developed by an elite group. </h4><blockquote><p>However, such approaches might be acceptable as a means of temporarily flagging potentially unethical situations until a representative sample of human ethical judgments can be obtained. </p><p>For example, a rule that said an AI can never provide information that might be used to harm other humans might flag potentially dangerous scenarios. </p><p>Where possible, responses to such situations could be delayed until they were reviewed by humans or otherwise subjected to deeper review. </p></blockquote><p>Some false positives will occur. Someone might ask about using arsenic to poison rats and have to wait for a response while the AI flags the question and gets other human agents to weigh in on whether answering it, given the context of the conversation, poses a risk to humans. As long as the delay is not too long, it might be acceptable if the delay prevents serious safety issues. Established mathematical methods for determining when a test is doing more harm than good, for example, in the medical profession, can be employed to help quantify these decisions. If we can use AI to determine in real time whether an applicant is a good credit risk, there is no reason that similar algorithms cannot be employed to delay or avoid answering certain potentially dangerous questions.</p><h4>Ideally, there would be a method for rapid review and appeal of the potentially dangerous cases. </h4><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BCh6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eee88e8-20f1-4125-8e23-e8fcdb9c6fc1_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BCh6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eee88e8-20f1-4125-8e23-e8fcdb9c6fc1_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!BCh6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eee88e8-20f1-4125-8e23-e8fcdb9c6fc1_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!BCh6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eee88e8-20f1-4125-8e23-e8fcdb9c6fc1_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!BCh6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eee88e8-20f1-4125-8e23-e8fcdb9c6fc1_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BCh6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eee88e8-20f1-4125-8e23-e8fcdb9c6fc1_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4eee88e8-20f1-4125-8e23-e8fcdb9c6fc1_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1404390,&quot;alt&quot;:&quot;Flowchart of an AI pausing an unfamiliar request for review by multiple AI agents, with human guidance resolving uncertain cases before a safe response.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/208291735?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eee88e8-20f1-4125-8e23-e8fcdb9c6fc1_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Flowchart of an AI pausing an unfamiliar request for review by multiple AI agents, with human guidance resolving uncertain cases before a safe response." title="Flowchart of an AI pausing an unfamiliar request for review by multiple AI agents, with human guidance resolving uncertain cases before a safe response." srcset="https://substackcdn.com/image/fetch/$s_!BCh6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eee88e8-20f1-4125-8e23-e8fcdb9c6fc1_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!BCh6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eee88e8-20f1-4125-8e23-e8fcdb9c6fc1_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!BCh6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eee88e8-20f1-4125-8e23-e8fcdb9c6fc1_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!BCh6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4eee88e8-20f1-4125-8e23-e8fcdb9c6fc1_1672x941.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Questionable requests are paused for fast review by multiple AI agents, and human judgment resolves cases that remain uncertain.</figcaption></figure></div><p>One approach, which minimizes delay to a fraction of a second while still providing some margin of safety, is to have questionable cases reviewed by multiple AI agents to see if there is consensus among them on the request&#8217;s safety. Human agents, working more slowly, could override the AI agents (teaching the AIs in the process) upon appeal or when they can get to the prioritized list of issues. Automated means for tracking the frequency and potential impact of unanticipated safety issues could help optimize the use of human decision-making for the most common and most important issues.</p><h4>These same techniques can improve accuracy even when no safety risk is involved. </h4><p>One preferred method for reducing hallucinations from LLMs is to have multiple AI agents process the same question and then take the consensus or majority answer as the most correct. This approach might employ versions of the same LLM with different parameter settings to generate multiple responses, or completely different LLM models. Users can set the degree of reliability they desire and are willing to pay for. That choice determines how much redundant processing is performed in the final answer.</p><p>The difference between employing imperfect real-time detection of safety issues using existing well-known approaches and doing nothing is huge. The nuances of balancing the opportunity costs of not responding to perfectly harmless questions with the costs of preventing disasters can be refined over time, ideally using a data-driven approach. However, real-time detection and prevention of issues before they occur is almost certainly a net positive, even at some threshold of false positives.</p><h4>All of this depends on AI acquiring human knowledge and values in the first place. </h4><p>From a user interface perspective, one of the simplest methods is for humans to have conversations with the AI they are customizing and then provide instructions to that AI on how it should behave when training other AIs. Such conversations can be initiated by either the AI being customized, the human doing the customization, or both. Humans with strong beliefs or knowledge about certain issues may want to focus conversations and subsequent AI customization in these areas.</p><p>In addition to conversing with humans, AI can conduct surveys to elicit their opinions and knowledge on a wide variety of subjects, including ethical views. Survey approaches have the advantage of being well-suited to gathering random, representative samples of human knowledge, using a variety of established online and offline survey methodologies. Like intelligent conversational approaches, where AI can guide the direction and content of the conversation to fill in knowledge gaps, survey methods can also target specific knowledge gaps, including gaps in coverage of certain ethical situations.</p><h4>Both conversational and survey methods require humans to engage with AI to teach it actively. </h4><p>However, humans have limited time to engage in such activities, and AI has an almost insatiable appetite for new knowledge. Therefore, AI will have to rely extensively on passive methods of knowledge acquisition, such as are currently employed in the creation of today&#8217;s LLMs. Any method that uses the passive digital footprints left by humans as they perform tasks, including online navigation, selecting products and websites, solving problems, and communicating with other humans, can be used to train AI and acquire knowledge.</p><p>Knowledge does not stand still. Some of what AI learns changes by the day, while the most fundamental human values change slowly, if at all. </p><p><em>The next post in this series examines the rate of change of knowledge, pictured as a spinning wheel with passing fashions out at the rim and the deepest human values at the motionless center.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/how-ai-can-catch-danger-in-real-time?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/how-ai-can-catch-danger-in-real-time?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/how-ai-can-catch-danger-in-real-time/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/how-ai-can-catch-danger-in-real-time/comments"><span>Leave a comment</span></a></p><p><em>This series draws on <a href="https://www.superintelligence.com/whitepaper-4-scalable-agi">White Paper 4: Safe, Scalable Artificial General Intelligence</a>. Read it in full to see how every piece fits together!</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/si-research-whitepapers&quot;,&quot;text&quot;:&quot;SuperIntelligence Design White Papers&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/si-research-whitepapers"><span>SuperIntelligence Design White Papers</span></a></p><div><hr></div><p><strong>If this made you think, subscribe to Superintelligence at <a href="http://read.superintelligence.com">read.superintelligence.com</a> so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading SuperIntelligence! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[Why Training AI Is Like Testing the Brakes]]></title><description><![CDATA[AI trained by millions of people learns the situations that matter most.]]></description><link>https://read.superintelligence.com/p/why-training-ai-is-like-testing-the</link><guid isPermaLink="false">https://read.superintelligence.com/p/why-training-ai-is-like-testing-the</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Fri, 31 Jul 2026 13:04:01 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!1ikp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63a8dac-823d-481b-bf1a-d88c1e2ee982_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1ikp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63a8dac-823d-481b-bf1a-d88c1e2ee982_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1ikp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63a8dac-823d-481b-bf1a-d88c1e2ee982_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!1ikp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63a8dac-823d-481b-bf1a-d88c1e2ee982_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!1ikp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63a8dac-823d-481b-bf1a-d88c1e2ee982_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!1ikp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63a8dac-823d-481b-bf1a-d88c1e2ee982_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1ikp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63a8dac-823d-481b-bf1a-d88c1e2ee982_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d63a8dac-823d-481b-bf1a-d88c1e2ee982_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1706288,&quot;alt&quot;:&quot;Editorial cover on a dark indigo background. On the left, dozens of blue scenario tiles are connected by blue lines to a large amber warning tile at the center, while several larger amber tiles around the edges also connect inward, emphasizing high-risk situations receiving focused attention. On the right, large cream text reads &#8220;Test What Matters Most,&#8221; with the subtitle &#8220;Safe AI learns the big risks first.&#8221; Below is the series branding: &#8220;SUPERINTELLIGENCE,&#8221; &#8220;Safe Scalable AGI Series,&#8221; and &#8220;by Dr. Craig A. Kaplan.&#8221;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/208278643?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63a8dac-823d-481b-bf1a-d88c1e2ee982_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Editorial cover on a dark indigo background. On the left, dozens of blue scenario tiles are connected by blue lines to a large amber warning tile at the center, while several larger amber tiles around the edges also connect inward, emphasizing high-risk situations receiving focused attention. On the right, large cream text reads &#8220;Test What Matters Most,&#8221; with the subtitle &#8220;Safe AI learns the big risks first.&#8221; Below is the series branding: &#8220;SUPERINTELLIGENCE,&#8221; &#8220;Safe Scalable AGI Series,&#8221; and &#8220;by Dr. Craig A. Kaplan.&#8221;" title="Editorial cover on a dark indigo background. On the left, dozens of blue scenario tiles are connected by blue lines to a large amber warning tile at the center, while several larger amber tiles around the edges also connect inward, emphasizing high-risk situations receiving focused attention. On the right, large cream text reads &#8220;Test What Matters Most,&#8221; with the subtitle &#8220;Safe AI learns the big risks first.&#8221; Below is the series branding: &#8220;SUPERINTELLIGENCE,&#8221; &#8220;Safe Scalable AGI Series,&#8221; and &#8220;by Dr. Craig A. Kaplan.&#8221;" srcset="https://substackcdn.com/image/fetch/$s_!1ikp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63a8dac-823d-481b-bf1a-d88c1e2ee982_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!1ikp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63a8dac-823d-481b-bf1a-d88c1e2ee982_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!1ikp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63a8dac-823d-481b-bf1a-d88c1e2ee982_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!1ikp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd63a8dac-823d-481b-bf1a-d88c1e2ee982_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4><span>If software were a car, we could live with an interior light failing more easily than we could live with the brakes failing. If we had to choose between testing interior lights and the brakes due to limited resources, we would prioritize testing the brakes. AI safety faces a similar problem. There are more dangerous situations than any training system could anticipate individually, so the practical goal is to cover the situations that arise most often and the ones where failure would cause the most harm.</span></h4><p><span>The design described in these white papers meets that goal with a large and diverse community of teachers. Millions of people each customize a personal AI agent, teaching it their knowledge and values. I call these agents Advanced Autonomous Artificial Intelligences, or AAAIs. Since people live in different circumstances and train their AAAIs based on those experiences, the community should collectively cover a broad range of ethical situations.</span></p><p><strong><span>Software engineers call this kind of problem path coverage, and at some level, the problem of training AI safely resembles it.</span></strong><span> </span></p><p><span>AI must be trained on enough representative situations involving dangerous or ethical decision-making so that its behavior becomes more reliable and trustworthy in those situations. Enough of the situations (paths) must be covered in the training. Current AI systems can hallucinate and behave unpredictably. This design aims to make AI substantially more predictable and trustworthy, especially in safety- and ethics-related contexts.</span></p><p><strong><span>Generally, when testing software, human developers create test cases to cover the use cases most likely to arise.</span></strong><span> </span></p><p><span>Since it is impossible to test every possible use of complex software, developers determine which use cases are most common and which have the highest impact if things go wrong. More dangerous scenarios get more testing than benign scenarios. Frequency matters too. If one interior light is used ten times as often as another, a failure in the more frequently used light would affect people ten times as often, so it deserves more of the limited testing. This logic applies when training AI, whether via RLHF, other AIs, or, as this design suggests, a combination of humans and many customized AI agents.</span></p><p><strong><span>Many representative humans customize AAAIs, and those humans and agents help train new AIs</span></strong><span>. </span></p><p><span>If the participating group is sufficiently broad and properly weighted, situations that occur frequently should also appear frequently in the training. In effect, the system samples both human values and the situations in which those values must be applied. The larger the sample, the more certain we can be that the most frequent cases have been addressed in ways aligned with the human population&#8217;s values.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!A_5J!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0346e3b9-11c1-420a-bb6c-9c935bb5ba60_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!A_5J!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0346e3b9-11c1-420a-bb6c-9c935bb5ba60_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!A_5J!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0346e3b9-11c1-420a-bb6c-9c935bb5ba60_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!A_5J!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0346e3b9-11c1-420a-bb6c-9c935bb5ba60_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!A_5J!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0346e3b9-11c1-420a-bb6c-9c935bb5ba60_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!A_5J!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0346e3b9-11c1-420a-bb6c-9c935bb5ba60_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0346e3b9-11c1-420a-bb6c-9c935bb5ba60_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1466636,&quot;alt&quot;:&quot;A dense cluster of blue dots labeled common situations beside a single amber dot connected to three specialists, labeled rare high-impact situations.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/208278643?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0346e3b9-11c1-420a-bb6c-9c935bb5ba60_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A dense cluster of blue dots labeled common situations beside a single amber dot connected to three specialists, labeled rare high-impact situations." title="A dense cluster of blue dots labeled common situations beside a single amber dot connected to three specialists, labeled rare high-impact situations." srcset="https://substackcdn.com/image/fetch/$s_!A_5J!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0346e3b9-11c1-420a-bb6c-9c935bb5ba60_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!A_5J!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0346e3b9-11c1-420a-bb6c-9c935bb5ba60_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!A_5J!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0346e3b9-11c1-420a-bb6c-9c935bb5ba60_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!A_5J!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0346e3b9-11c1-420a-bb6c-9c935bb5ba60_1672x941.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Common situations are covered by density; rare dangers are assigned deliberately.</figcaption></figure></div><blockquote><p><strong><span>Addressing some dangerous cases and difficult ethical decisions is more challenging</span></strong><span>. </span></p><p><strong><span>That&#8217;s because dangerous situations and tough ethical decisions are often relatively rare. In this case, the approach is to ask humans and AI agents to think of as many dangerous scenarios as possible. The total pool can then be allocated among the human and AI samples so that identified high-impact scenarios receive input from enough different agents to provide a representative range of judgments. Input can then be directed toward people with relevant experience.</span></strong></p></blockquote><p><span>Suppose an AI must decide which patients receive medical attention during triage. Emergency room doctors and paramedics, who are used to making triage decisions, may recognize that devoting scarce resources to a patient with almost no chance of survival could reduce the chances of saving other patients. A well-meaning person without triage experience may understandably find that choice harder to make. For this difficult ethical situation, specialized knowledge is an advantage, and we might prefer to let the medically experienced professionals teach the AI. People also tend to identify dilemmas from the worlds they know. Emergency clinicians will think of triage, while an HR professional might contribute scenarios involving hiring, promotion, and fairness. A large and varied population, therefore, generates a broader set of ethical cases than a small, centrally selected group.</span></p><p><span>A broad community does two things at once: it generates a wider range of ethical situations and connects them with people who understand them. In ordinary life, people gain the most experience with choices they face repeatedly and give special attention to decisions with serious consequences. A community of teachers can create the same pattern in AI training.</span></p><p><span>Professional human trainers can use RLHF to address gaps in handling important but infrequent ethical dilemmas. This gives the student AI broader and more representative ethical training. Of course, no amount of advance training can anticipate everything. </span></p><p><em><span>The next post examines how AI can detect danger in real time, flag situations it cannot judge, and pause for human guidance.</span></em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/why-training-ai-is-like-testing-the?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/why-training-ai-is-like-testing-the?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/why-training-ai-is-like-testing-the/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/why-training-ai-is-like-testing-the/comments"><span>Leave a comment</span></a></p><div><hr></div><p><strong><span>This series draws on </span></strong><em><strong><a href="https://www.superintelligence.com/whitepaper-4-scalable-agi"><span>White Paper 4: Safe, Scalable Artificial General Intelligence</span></a></strong></em><strong><span>. Read it in full to see how every piece fits together!</span></strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper-4-scalable-agi&quot;,&quot;text&quot;:&quot;SuperIntelligence Design White Papers&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper-4-scalable-agi"><span>SuperIntelligence Design White Papers</span></a></p><p><strong><span>If this made you think, subscribe to Superintelligence at </span><a href="https://read.superintelligence.com/"><span>read.superintelligence.com </span></a><span>so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</span></strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading SuperIntelligence! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[AI and the Statistics of Right and Wrong]]></title><description><![CDATA[Diversified human values protect AI safety the way a diversified portfolio protects an investor.]]></description><link>https://read.superintelligence.com/p/ai-and-the-statistics-of-right-and</link><guid isPermaLink="false">https://read.superintelligence.com/p/ai-and-the-statistics-of-right-and</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Wed, 29 Jul 2026 13:03:58 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!RjOO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f5a15e-1722-41dc-8245-f0ffc5d9d877_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!RjOO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f5a15e-1722-41dc-8245-f0ffc5d9d877_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!RjOO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f5a15e-1722-41dc-8245-f0ffc5d9d877_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!RjOO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f5a15e-1722-41dc-8245-f0ffc5d9d877_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!RjOO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f5a15e-1722-41dc-8245-f0ffc5d9d877_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!RjOO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f5a15e-1722-41dc-8245-f0ffc5d9d877_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!RjOO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f5a15e-1722-41dc-8245-f0ffc5d9d877_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a4f5a15e-1722-41dc-8245-f0ffc5d9d877_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1827770,&quot;alt&quot;:&quot;many hands supporting a amber globe, Safety in Numbers, Many values make AI safer than one. Superintelligence Safe Scalable AI series by Dr. Craig A. Kaplan&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/208010026?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f5a15e-1722-41dc-8245-f0ffc5d9d877_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="many hands supporting a amber globe, Safety in Numbers, Many values make AI safer than one. Superintelligence Safe Scalable AI series by Dr. Craig A. Kaplan" title="many hands supporting a amber globe, Safety in Numbers, Many values make AI safer than one. Superintelligence Safe Scalable AI series by Dr. Craig A. Kaplan" srcset="https://substackcdn.com/image/fetch/$s_!RjOO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f5a15e-1722-41dc-8245-f0ffc5d9d877_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!RjOO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f5a15e-1722-41dc-8245-f0ffc5d9d877_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!RjOO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f5a15e-1722-41dc-8245-f0ffc5d9d877_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!RjOO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa4f5a15e-1722-41dc-8245-f0ffc5d9d877_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4><span>Although it is undesirable to have the safety and ethics of AI driven by a constitution written by a small, elite group, it is still possible that most humans, regardless of cultural or individual differences, would agree on certain normative ethical principles. </span></h4><p><strong><span>Examples of these might include variants of:</span></strong></p><ul><li><p><strong><span>The Golden Rule (Do not do to others that which you would not like done to you)</span></strong></p></li><li><p><strong><span>First, do no harm (as reflected in the Hippocratic Oath taken by physicians)</span></strong></p></li><li><p><strong><span>Do not kill (unnecessarily or except in specific, exceptional circumstances)</span></strong></p></li><li><p><strong><span>Preserve individual freedom (unless it limits the freedom of others)</span></strong></p></li></ul><p><span>Of course, almost as soon as you read these principles, exceptions spring to mind. Do not kill, but what about self-defense or war? Preserve individual freedom, but what are the limits, and when does it impinge on others? Details and nuance matter, even when applying principles that most humans would broadly embrace. </span></p><blockquote><h4><span>However, by starting with general normative ethics widely accepted by a large, diverse, and representative group of humans, it is possible to refine these principles and determine when and how they apply in detailed circumstances much more efficiently than if no starting principles existed at all. General ethical norms are a point of departure that can help AI achieve realistic, nuanced ethics and behavior aligned with what most humans believe is good and aspire to.</span></h4></blockquote><p><span>Once we have admitted the potential usefulness of ethical norms as a starting point for further refinement, the door is open to group and planetary norms. </span></p><p><span>There is a continuum: at one end are highly individualized AIs trained to think and act like a particular person; farther along are AIs trained to reflect the values of specific groups; and at the far end are AIs guided by ethical norms shared across many groups. Ethical norms at each point on the continuum can serve as a starting point for training AI to exhibit ethical behavior. The idea that one set of norms or one constitution should power all of AI is likely unrealistic and far too brittle to work in the real world. </span></p><p><span>If it were possible, then the many differing viewpoints espoused by religious, political, and cultural groups would long ago have merged into a consensus. The diversity in human ethical norms is not a bug; it is a feature. We should not expect AI to achieve consensus and maintain human alignment if humans themselves cannot, especially if there is debate over whether such a consensus is even desirable.</span></p><h4><span>Humans also often enter into ethical, implicit, or explicit social contracts when they join a group or participate in society. </span></h4><p><span>Members of a particular religion largely agree with a set of rules and ethical precepts espoused by that religion, often enshrined in holy books. Similarly, Confucianism in China, the ideals reflected in the Declaration of Independence and Constitution in the USA, and liberal or conservative ideologies for various political groups all contain normative prescriptions for human behavior. By being a citizen of, or simply living in, a particular country, humans are explicitly subject to the laws of that country, including laws that explicitly specify what criminal (aka wrong) behavior is. Thus, for ethical problems, the solution sometimes depends on what social or ethical contract humans have made with the group or culture in which they find themselves. Such contracts can be useful for simplifying AI training, since a starting point can be the laws of a particular country or the implicit or explicit rules of a particular group.</span></p><h4><span>It has been said that democracy is a bad political system, but that all the others are worse. </span></h4><p><span>Most humans would agree that the most important concern regarding AI is the existential threat it currently poses to the majority. That is, AI could wipe humans out. If that happens, it doesn&#8217;t matter what form of government or religion you prefer. We&#8217;d all be dead, and the point would be moot. So we should be asking not which religion or form of government is best, but which principles are most likely to lead to humanity&#8217;s survival.</span></p><h4><span>Democratically representing the opinions of most people is rarely optimal, but generally achieves an acceptable outcome. </span></h4><p><span>Collective intelligence (the idea that two heads are better than one) is responsible for the vast majority of human progress, culture, and technology. But when it comes to the subjective area of ethics and values, where there is no objectively correct answer, just human opinions, democracy, or a collective intelligence approach, if you prefer, really shines. One benefit of a democratic and representative set of human values is that it tends to mitigate extreme positions, which are likely to pose the greatest risk to human survival. There is a beneficial diversification effect regarding values.</span></p><p><span>Just as diversification in an asset portfolio reduces volatility and risk, so too does a diversity of human opinions and judgments tend to stabilize the overall portfolio of values. In an asset portfolio, a diversified portfolio always returns less than if you were to concentrate all the investment on the top winners. The problem is that no one knows who the winners will be with any certainty. That is why the diversified approach of just buying the index tends to outperform more than 80 percent of all portfolio managers who try to beat the market.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!tgIY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98543b10-c5f3-47e4-a3e1-4e286c5b8827_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!tgIY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98543b10-c5f3-47e4-a3e1-4e286c5b8827_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!tgIY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98543b10-c5f3-47e4-a3e1-4e286c5b8827_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!tgIY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98543b10-c5f3-47e4-a3e1-4e286c5b8827_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!tgIY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98543b10-c5f3-47e4-a3e1-4e286c5b8827_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!tgIY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98543b10-c5f3-47e4-a3e1-4e286c5b8827_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/98543b10-c5f3-47e4-a3e1-4e286c5b8827_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1481466,&quot;alt&quot;:&quot;Infographic comparing a leaning tower labeled a few voices with a wide, stable brick pyramid labeled millions of voices, topped by a glowing amber sphere.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/208010026?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98543b10-c5f3-47e4-a3e1-4e286c5b8827_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Infographic comparing a leaning tower labeled a few voices with a wide, stable brick pyramid labeled millions of voices, topped by a glowing amber sphere." title="Infographic comparing a leaning tower labeled a few voices with a wide, stable brick pyramid labeled millions of voices, topped by a glowing amber sphere." srcset="https://substackcdn.com/image/fetch/$s_!tgIY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98543b10-c5f3-47e4-a3e1-4e286c5b8827_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!tgIY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98543b10-c5f3-47e4-a3e1-4e286c5b8827_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!tgIY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98543b10-c5f3-47e4-a3e1-4e286c5b8827_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!tgIY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F98543b10-c5f3-47e4-a3e1-4e286c5b8827_1672x941.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Similarly, there are philosopher kings or religious saints who can make laws and ethical rules that, for a time, are far superior to the collective judgment and behavior of the masses. But what happens when the superior king or saint is gone? Then a power-hungry dictator might arise whose reign is far worse. The more stable approach (less likely to be really great, but also less likely to be really terrible) is to follow the values of a large representative population of humans. All these humans want to survive. Most want good things for themselves and their fellow humans. Few want to destroy the environment or the planet. While the collective values are imperfect, they are usually not malevolent. Importantly, they are based on human hearts!</span></p><p><span>Putting aside the practical benefits of a diversified, representative portfolio approach to human values, a representative sample is also a scientifically valid way to accurately answer a question with no logical answer: what is right and what is wrong according to humans. A fast computer could answer a math problem faster than a million humans, but when it comes to the subjective determination of what is wrong and what is right, calculation speed is useless. If we want to know what human values are, there is no substitute for asking them and watching their behavior. The more humans we ask and watch, the more representative the values may be.</span></p><p><span>Of course, there are potentially an infinite number of dangerous situations we need to train AI to handle safely, and training resources are limited. The next post examines path coverage, or how to decide which situations matter most, and why, if software were a car, we would test the brakes before the interior lights.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share&quot;,&quot;text&quot;:&quot;Share SuperIntelligence&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/?utm_source=substack&amp;utm_medium=email&amp;utm_content=share&amp;action=share"><span>Share SuperIntelligence</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/ai-and-the-statistics-of-right-and/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/ai-and-the-statistics-of-right-and/comments"><span>Leave a comment</span></a></p><div><hr></div><p><strong><span>This series draws on </span></strong><em><strong><a href="https://www.superintelligence.com/whitepaper-4-scalable-agi"><span>White Paper 4: Safe, Scalable Artificial General Intelligence</span></a></strong></em><strong><span>. Read it in full to see how every piece fits together!</span></strong></p><p><strong><span>If this made you think, subscribe to Superintelligence at read.superintelligence.com so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</span></strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/si-research-whitepapers&quot;,&quot;text&quot;:&quot;SuperIntelligence Design White Papers&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/si-research-whitepapers"><span>SuperIntelligence Design White Papers</span></a></p>]]></content:encoded></item><item><title><![CDATA[How AI Should Decide When There Is No Right Answer]]></title><description><![CDATA[The Trolley Problem reveals the human values that no law can write down, and AI can learn them.]]></description><link>https://read.superintelligence.com/p/how-ai-should-decide-when-there-is</link><guid isPermaLink="false">https://read.superintelligence.com/p/how-ai-should-decide-when-there-is</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Mon, 27 Jul 2026 12:49:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!7ARn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe3a70c5-62bb-42a2-b3a2-1015b701a8b4_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7ARn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe3a70c5-62bb-42a2-b3a2-1015b701a8b4_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7ARn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe3a70c5-62bb-42a2-b3a2-1015b701a8b4_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!7ARn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe3a70c5-62bb-42a2-b3a2-1015b701a8b4_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!7ARn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe3a70c5-62bb-42a2-b3a2-1015b701a8b4_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!7ARn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe3a70c5-62bb-42a2-b3a2-1015b701a8b4_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7ARn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe3a70c5-62bb-42a2-b3a2-1015b701a8b4_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/be3a70c5-62bb-42a2-b3a2-1015b701a8b4_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1978999,&quot;alt&quot;:&quot;A level balance scale made of glowing blue particles weighing two groups of human figures, with a single amber point of light at its center, on a dark field.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/208008528?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe3a70c5-62bb-42a2-b3a2-1015b701a8b4_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A level balance scale made of glowing blue particles weighing two groups of human figures, with a single amber point of light at its center, on a dark field." title="A level balance scale made of glowing blue particles weighing two groups of human figures, with a single amber point of light at its center, on a dark field." srcset="https://substackcdn.com/image/fetch/$s_!7ARn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe3a70c5-62bb-42a2-b3a2-1015b701a8b4_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!7ARn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe3a70c5-62bb-42a2-b3a2-1015b701a8b4_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!7ARn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe3a70c5-62bb-42a2-b3a2-1015b701a8b4_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!7ARn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbe3a70c5-62bb-42a2-b3a2-1015b701a8b4_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4><span>Some of the hardest ethical decisions have no right answer, and AI must act anyway. </span></h4><h4><span>The moral sense and opinions of individual humans regulate the vast majority of human behavior. And there is a huge range of decisions people make every day for which there is no right moral answer, just opinions about what is right or wrong. </span></h4><h4><span>Given this complex situation and the fact that machines are notorious for requiring exact specifications to behave, implementing ethics-specific AI safety solutions may differ from the methods that work for other types of knowledge.</span></h4><p><span>Consider a classic example of an ethical dilemma well-known in the field of AI ethics, the </span><a href="https://en.wikipedia.org/wiki/Trolley_problem"><span>Trolley Problem</span></a><span>. In one version of this dilemma, a self-driving car controlled by an AI finds itself having to choose between killing pedestrians who suddenly jump in front of the car or swerving into a barrier to avoid the pedestrians and killing the occupants of the car. Like many difficult ethical decisions, there is no right answer. Yet humans still have opinions about what is ethical and what they would do in such a situation.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6esm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f60f2c-d6df-41e4-9e55-a2c494014226_1586x992.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6esm!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f60f2c-d6df-41e4-9e55-a2c494014226_1586x992.png 424w, https://substackcdn.com/image/fetch/$s_!6esm!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f60f2c-d6df-41e4-9e55-a2c494014226_1586x992.png 848w, https://substackcdn.com/image/fetch/$s_!6esm!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f60f2c-d6df-41e4-9e55-a2c494014226_1586x992.png 1272w, https://substackcdn.com/image/fetch/$s_!6esm!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f60f2c-d6df-41e4-9e55-a2c494014226_1586x992.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6esm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f60f2c-d6df-41e4-9e55-a2c494014226_1586x992.png" width="1456" height="911" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/91f60f2c-d6df-41e4-9e55-a2c494014226_1586x992.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:911,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2013034,&quot;alt&quot;:&quot;nfographic titled &#8220;The Trolley Problem&#8221; showing a self-driving car approaching a fork in the road. On the left, pedestrians are in a crosswalk with the option &#8220;Continue Forward: Pedestrians at Risk.&#8221; On the right, concrete barriers represent the option &#8220;Swerve: Passengers at Risk.&#8221; An AI symbol above the car weighs contextual factors, including whether pedestrians are crossing legally, who is in the road, who is in the car, and their ages and circumstances. Bottom text reads: &#8220;There Is No Right Answer. Humans Weigh Context &#8212; AI Must Decide How.&#8221;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/208008528?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f60f2c-d6df-41e4-9e55-a2c494014226_1586x992.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="nfographic titled &#8220;The Trolley Problem&#8221; showing a self-driving car approaching a fork in the road. On the left, pedestrians are in a crosswalk with the option &#8220;Continue Forward: Pedestrians at Risk.&#8221; On the right, concrete barriers represent the option &#8220;Swerve: Passengers at Risk.&#8221; An AI symbol above the car weighs contextual factors, including whether pedestrians are crossing legally, who is in the road, who is in the car, and their ages and circumstances. Bottom text reads: &#8220;There Is No Right Answer. Humans Weigh Context &#8212; AI Must Decide How.&#8221;" title="nfographic titled &#8220;The Trolley Problem&#8221; showing a self-driving car approaching a fork in the road. On the left, pedestrians are in a crosswalk with the option &#8220;Continue Forward: Pedestrians at Risk.&#8221; On the right, concrete barriers represent the option &#8220;Swerve: Passengers at Risk.&#8221; An AI symbol above the car weighs contextual factors, including whether pedestrians are crossing legally, who is in the road, who is in the car, and their ages and circumstances. Bottom text reads: &#8220;There Is No Right Answer. Humans Weigh Context &#8212; AI Must Decide How.&#8221;" srcset="https://substackcdn.com/image/fetch/$s_!6esm!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f60f2c-d6df-41e4-9e55-a2c494014226_1586x992.png 424w, https://substackcdn.com/image/fetch/$s_!6esm!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f60f2c-d6df-41e4-9e55-a2c494014226_1586x992.png 848w, https://substackcdn.com/image/fetch/$s_!6esm!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f60f2c-d6df-41e4-9e55-a2c494014226_1586x992.png 1272w, https://substackcdn.com/image/fetch/$s_!6esm!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F91f60f2c-d6df-41e4-9e55-a2c494014226_1586x992.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Surveys of many people have shown that what humans consider ethical depends. It depends on who is in the car, who the pedestrians are, whether they are crossing legally or illegally, and even how old the people involved are. Humans are more likely to instruct the car to run over pedestrians if the pedestrians are crossing illegally, are homeless, or are simply old. According to the survey research, humans are less likely to kill, or allow to be killed, those who are young, pregnant, women, or in certain professions, such as the medical profession.</span></p><p><span>None of these aspects of human decision-making is captured in our legal system. No law says it is okay to run over someone older than someone younger, yet humans consider these factors, and mathematical weights can be assigned to each. Similarly, AI can learn to make decisions, taking these same weights into account.</span></p><p><span>To have AIs that behave in ways that make sense to most humans, AI will have to be trained not according to a rigid constitution but rather according to how real humans actually behave. That behavior varies across cultures. In the US, something of a youth culture, running over elderly pedestrians is likely more acceptable than in certain Asian cultures, where elders are revered and held in high esteem. If AI is to make ethical decisions in the same way that most humans do, it will have to take these cultural factors into account.</span></p><h4><span>Capturing that diversity is where the standard approaches come up short. </span></h4><p><span>Constitutional AI falls short because we would need a different constitution for each human group. Having AI interact with many humans to learn their values, as in RLHF, would be much more effective, RLHF scales poorly. The best way to capture the wide diversity of human values while retaining the scalability that comes with AI involved in instruction is to have a multitude of teachers, both human and AI agents. Each AI agent should be trained by a different person, so that it carries the unique values, ethics, and moral sensibility of its owner into every interaction, including interactions with other AIs.</span></p><h4><span>In the long run, AI will undoubtedly surpass human ability in cognition, problem-solving, and information processing. </span></h4><p><span>As AI grows increasingly intelligent and capable, the role of humans will increasingly be to determine the values and fundamental goals that the more intelligent AIs seek to realize! That role should not belong to a small elite group of programmers; instead, it should reflect as broad a cross-section of humanity as possible. </span></p><p><span>By including all humans who are able and willing to customize their AIs in the crucial task of determining AI-based values, we can achieve broad representation more efficiently and cost-effectively than any existing approach. Humans can, and should, remain in the loop as much as possible when training AI. To the degree that humans are unavailable, or the resource demands are too great for all the training to be done by humans themselves, the next best thing is to include a wide and diverse group of AI agents in the training, each customized by a different person to reflect that person&#8217;s values and ethics. Once a human trains an AI agent, it can operate around the clock, with or without its original owner&#8217;s supervision, allowing the owner&#8217;s values to shape training and other activities without requiring constant human involvement.</span></p><p><span>Of course, even without a right answer, it is still possible that most humans, regardless of cultural or individual differences, would agree on certain normative ethical principles. The next post examines those ethical norms and the statistics of right and wrong, including why a broad portfolio of human values protects against moral catastrophe, just as a diversified portfolio protects an investor.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/how-ai-should-decide-when-there-is?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/how-ai-should-decide-when-there-is?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/how-ai-should-decide-when-there-is/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/how-ai-should-decide-when-there-is/comments"><span>Leave a comment</span></a></p><div class="directMessage button" data-attrs="{&quot;userId&quot;:163471793,&quot;userName&quot;:&quot;Dr. Craig A. Kaplan&quot;,&quot;canDm&quot;:null,&quot;dmUpgradeOptions&quot;:null,&quot;isEditorNode&quot;:true}" data-component-name="DirectMessageToDOM"></div><div><hr></div><p><em><span>This series draws on </span><a href="https://www.superintelligence.com/whitepaper-4-scalable-agi"><span>White Paper 4: Safe, Scalable Artificial General Intelligence</span></a><span>. Read it in full to see how every piece fits together!</span></em></p><p><strong><span>If this made you think, subscribe to Superintelligence at read.superintelligence.com so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</span></strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/si-research-whitepapers&quot;,&quot;text&quot;:&quot;SuperIntelligence Design White Papers&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/si-research-whitepapers"><span>SuperIntelligence Design White Papers</span></a></p>]]></content:encoded></item><item><title><![CDATA[Who Should Have the Most Influence on AI?]]></title><description><![CDATA[Two methods can combine the values of many, but the harder question is how much each voice should count.]]></description><link>https://read.superintelligence.com/p/who-should-have-the-most-influence</link><guid isPermaLink="false">https://read.superintelligence.com/p/who-should-have-the-most-influence</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Fri, 24 Jul 2026 12:46:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!COYW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3242eda0-262c-4ca1-9e0d-bf62e845c8ac_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!COYW!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3242eda0-262c-4ca1-9e0d-bf62e845c8ac_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!COYW!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3242eda0-262c-4ca1-9e0d-bf62e845c8ac_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!COYW!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3242eda0-262c-4ca1-9e0d-bf62e845c8ac_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!COYW!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3242eda0-262c-4ca1-9e0d-bf62e845c8ac_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!COYW!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3242eda0-262c-4ca1-9e0d-bf62e845c8ac_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!COYW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3242eda0-262c-4ca1-9e0d-bf62e845c8ac_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3242eda0-262c-4ca1-9e0d-bf62e845c8ac_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1624887,&quot;alt&quot;:&quot;Line illustration of a patient deciding between a credentialed surgeon and a trusted friend, with amber bars showing the friend's advice weighing more. Text: Trust Can Outweigh Expertise. A trusted friend's judgment can count for more than a wall of credentials. Superintelligence, Safe Scalable AGI Series, by Dr. Craig A. Kaplan.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/207873900?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3242eda0-262c-4ca1-9e0d-bf62e845c8ac_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Line illustration of a patient deciding between a credentialed surgeon and a trusted friend, with amber bars showing the friend's advice weighing more. Text: Trust Can Outweigh Expertise. A trusted friend's judgment can count for more than a wall of credentials. Superintelligence, Safe Scalable AGI Series, by Dr. Craig A. Kaplan." title="Line illustration of a patient deciding between a credentialed surgeon and a trusted friend, with amber bars showing the friend's advice weighing more. Text: Trust Can Outweigh Expertise. A trusted friend's judgment can count for more than a wall of credentials. Superintelligence, Safe Scalable AGI Series, by Dr. Craig A. Kaplan." srcset="https://substackcdn.com/image/fetch/$s_!COYW!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3242eda0-262c-4ca1-9e0d-bf62e845c8ac_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!COYW!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3242eda0-262c-4ca1-9e0d-bf62e845c8ac_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!COYW!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3242eda0-262c-4ca1-9e0d-bf62e845c8ac_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!COYW!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3242eda0-262c-4ca1-9e0d-bf62e845c8ac_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4><span>Imagine millions of individuals, each customizing a personal AI agent and teaching it their knowledge and their values. </span></h4><h4><span>Those humans and their agents together shape the AIs that come after them, so the ethics of any new AI become the property of a whole community. </span></h4><h4><span>Turning that community into a single trained student is a practical matter of mechanics, namely, how the lessons of millions of teachers become the values of one AI. </span></h4><h4><span>There are two roads, and either one can get there.</span></h4><h4><span>The first road runs through feedback. </span></h4><p><span>An LLM being trained adjusts its network weights based on the responses it receives, much as it does during Reinforcement Learning with Human Feedback. The difference is that the feedback can come from both humans and AI agents. This is the Reinforcement Learning via Feedback, or RLF, introduced in the </span><a href="https://open.substack.com/pub/superintelligencebyiq/p/who-should-teach-ai-right-from-wrong"><span>previous post</span></a><span>. Because AI agents can provide feedback much faster and at far lower cost than humans, a large community of customized agents could evaluate the same scenario together and help train a new AI.</span></p><h4><span>The second road proposes combining weights directly.</span></h4><p><span>Imagine two copies of the same LLM, one customized through interactions with one person and the other customized by someone else. Training has changed each model&#8217;s network weights, leaving two distinct models that carry the influence of two different people. The weights of those models could then be combined mathematically to create a third model that reflects input from both teachers. In the simplest proposed scheme, averaging the weights would give the two teachers equal influence, whereas other combinations could give one teacher more influence than the other.</span></p><h4><span>For the purpose of comparing weighting schemes, the two roads are functionally equivalent. </span></h4><p><span>Many teachers can shape a single student through feedback, or separately trained models can contribute through direct weight combination, and the choice of path may depend on factors such as the availability of teachers and the computational resources at hand. Either road leads to the more important question of how much each teacher should count.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!b1ge!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29d89499-d935-4994-8915-2ded3fb4cd95_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!b1ge!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29d89499-d935-4994-8915-2ded3fb4cd95_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!b1ge!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29d89499-d935-4994-8915-2ded3fb4cd95_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!b1ge!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29d89499-d935-4994-8915-2ded3fb4cd95_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!b1ge!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29d89499-d935-4994-8915-2ded3fb4cd95_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!b1ge!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29d89499-d935-4994-8915-2ded3fb4cd95_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/29d89499-d935-4994-8915-2ded3fb4cd95_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1512094,&quot;alt&quot;:&quot;Wide editorial diagram, 1536 by 1024 landscape, flat clean style, thick simple line work, large type throughout, no fine detail. Deep ink-blue background, hex 1d3557, no grid. Headline at top in large off-white bold sans-serif capitals: TWO ROADS TO ONE STUDENT AI. Below, two horizontal rows flowing left to right, both ending at one large glowing circle at the right center of the frame, blended cool blue, hex 4a90d9, and warm amber, hex e8a13d, labeled inside in large off-white letters: STUDENT AI. Top row, labeled ROAD ONE FEEDBACK: two large icons side by side, one cool-blue human silhouette and one amber robot head, joined by thick arrows into one large rounded rectangle labeled MODEL, then one thick arrow to the student circle. Bottom row, labeled ROAD TWO COMBINING WEIGHTS: two large rounded rectangles each labeled MODEL, thick arrows into a large circle with a bold plus sign, then one thick arrow to the student circle. One short line at the bottom in large off-white sans-serif: Both roads reach the same student. No terminal periods on labels, generous spacing, amber for AI, cool blue for humans, everything sized to stay readable at phone width.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/207873900?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29d89499-d935-4994-8915-2ded3fb4cd95_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Wide editorial diagram, 1536 by 1024 landscape, flat clean style, thick simple line work, large type throughout, no fine detail. Deep ink-blue background, hex 1d3557, no grid. Headline at top in large off-white bold sans-serif capitals: TWO ROADS TO ONE STUDENT AI. Below, two horizontal rows flowing left to right, both ending at one large glowing circle at the right center of the frame, blended cool blue, hex 4a90d9, and warm amber, hex e8a13d, labeled inside in large off-white letters: STUDENT AI. Top row, labeled ROAD ONE FEEDBACK: two large icons side by side, one cool-blue human silhouette and one amber robot head, joined by thick arrows into one large rounded rectangle labeled MODEL, then one thick arrow to the student circle. Bottom row, labeled ROAD TWO COMBINING WEIGHTS: two large rounded rectangles each labeled MODEL, thick arrows into a large circle with a bold plus sign, then one thick arrow to the student circle. One short line at the bottom in large off-white sans-serif: Both roads reach the same student. No terminal periods on labels, generous spacing, amber for AI, cool blue for humans, everything sized to stay readable at phone width." title="Wide editorial diagram, 1536 by 1024 landscape, flat clean style, thick simple line work, large type throughout, no fine detail. Deep ink-blue background, hex 1d3557, no grid. Headline at top in large off-white bold sans-serif capitals: TWO ROADS TO ONE STUDENT AI. Below, two horizontal rows flowing left to right, both ending at one large glowing circle at the right center of the frame, blended cool blue, hex 4a90d9, and warm amber, hex e8a13d, labeled inside in large off-white letters: STUDENT AI. Top row, labeled ROAD ONE FEEDBACK: two large icons side by side, one cool-blue human silhouette and one amber robot head, joined by thick arrows into one large rounded rectangle labeled MODEL, then one thick arrow to the student circle. Bottom row, labeled ROAD TWO COMBINING WEIGHTS: two large rounded rectangles each labeled MODEL, thick arrows into a large circle with a bold plus sign, then one thick arrow to the student circle. One short line at the bottom in large off-white sans-serif: Both roads reach the same student. No terminal periods on labels, generous spacing, amber for AI, cool blue for humans, everything sized to stay readable at phone width." srcset="https://substackcdn.com/image/fetch/$s_!b1ge!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29d89499-d935-4994-8915-2ded3fb4cd95_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!b1ge!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29d89499-d935-4994-8915-2ded3fb4cd95_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!b1ge!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29d89499-d935-4994-8915-2ded3fb4cd95_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!b1ge!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F29d89499-d935-4994-8915-2ded3fb4cd95_1536x1024.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Feedback and weight combination are two roads to the same student, which makes the question of how much each teacher counts a real design choice.</figcaption></figure></div><h4><span>The natural starting point is one agent, one vote. </span></h4><p><span>Each participating human or customized AI receives equal influence over the student, and only one copy of each customized AI participates, so every teacher counts exactly once. However, equal influence is only one possible choice.</span></p><h4><span>A first variation gives different weights to human and AI input. </span></h4><p><span>One might expect humans to count more because people are usually better equipped to represent their own values than the AI agents they trained, yet there can be exceptions. Imagine that someone spends years teaching a personal AI agent their values and preferences, and later develops a serious cognitive impairment. The agent might then represent the person&#8217;s long-established values more faithfully than the person can express them, and in such a case, it might make sense for the agent&#8217;s input to carry more weight.</span></p><p><span>That exception must not become a reason for AI to dismiss human input altogether. Human values remain fundamental, even when AI can find better, faster, and more effective ways to pursue human goals. Later in this series, we return to the question of why greater intelligence does not give AI the authority to determine humanity&#8217;s values.</span></p><h4><span>Another possible scheme gives more weight to expertise. </span></h4><p><span>Physicians who have spent their careers advising terminally ill patients may have more insight into end-of-life decisions than someone who has never faced them. The design does not insist that experts should always count more; it allows input to be weighted by expertise and by its relevance to the decision being made. People are often reluctant to surrender important personal decisions to professional experts, and prevailing values preserve the &#8220;right&#8221; of people to do stupid things, as long as they are not harming others.</span></p><blockquote><p><strong><span>Expertise is also different from trust. A patient deciding whether to have an operation might seek advice from a highly knowledgeable surgeon. If that surgeon has a reputation for recommending unnecessary procedures, the patient might place greater trust in a family friend with good judgment and the patient&#8217;s best interests at heart. Of course, an expert who is also highly trusted might deserve the greatest influence of all. AI can weigh its teachers in similar ways. </span></strong></p><p><strong><span>Each human or AI agent can carry metadata describing relevant expertise, trustworthiness, background, preferences, and past performance, and that metadata can adjust the influence given to the agent&#8217;s input. The same approach can apply beyond values, helping to combine knowledge, skills, and judgment from many different sources.</span></strong></p></blockquote><p><span>Time can affect weighting as well. Recent input may better reflect current circumstances, and ethical norms, institutional policies, and laws change. In 2020, Disney placed warnings before some older films stating that they contained negative depictions or mistreatment of people or cultures, and that those stereotypes were wrong. By 2025, the company was using shorter language, stating that the program was presented as originally created and might contain negative depictions. The revisions show that institutions can change how they characterize the same content within a few years. </span></p><blockquote><p><strong><span>Sometimes change is gradual, and older input can slowly fade in influence, while at other times change is immediate. </span></strong></p><p><strong><span>During Prohibition, selling alcohol was illegal, and when Prohibition ended, its legal status changed overnight. A system that uses law as one signal of prevailing norms would need a stepwise adjustment, since no gradual fade can capture a change like that.</span></strong></p></blockquote><p><span>These weighting schemes explain how the values of many teachers can be pooled. They do not tell us what to do when sincere and thoughtful people give different answers, and no objectively correct answer exists. Researchers have documented how differently people around the world resolve one famous dilemma about who should be spared and who should be sacrificed. </span></p><p><span>My next post takes up that dilemma, the trolley problem, and shows why disagreement strengthens the case for teaching AI the values of many people instead of a chosen few.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/who-should-have-the-most-influence?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/who-should-have-the-most-influence?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/who-should-have-the-most-influence/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/who-should-have-the-most-influence/comments"><span>Leave a comment</span></a></p><div><hr></div><blockquote><p>This series draws on <a href="https://www.superintelligence.com/whitepaper-4-scalable-agi">White Paper 4: Safe, Scalable Artificial General Intelligence</a>. Read it in full to see how every piece fits together!</p><p><em><strong>If this made you think, subscribe to Superintelligence at <a href="https://read.superintelligence.com/">read.superintelligence.com</a> so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</strong></em></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/si-research-whitepapers&quot;,&quot;text&quot;:&quot;SuperIntelligence Design White Papers&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/si-research-whitepapers"><span>SuperIntelligence Design White Papers</span></a></p>]]></content:encoded></item><item><title><![CDATA[Who Should Teach AI Right from Wrong?]]></title><description><![CDATA[AI learns better values when more people help teach it.]]></description><link>https://read.superintelligence.com/p/who-should-teach-ai-right-from-wrong</link><guid isPermaLink="false">https://read.superintelligence.com/p/who-should-teach-ai-right-from-wrong</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Wed, 22 Jul 2026 12:45:15 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!amev!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd658f855-99fa-438d-88a0-f13bba2c1b69_1669x942.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!amev!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd658f855-99fa-438d-88a0-f13bba2c1b69_1669x942.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!amev!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd658f855-99fa-438d-88a0-f13bba2c1b69_1669x942.png 424w, https://substackcdn.com/image/fetch/$s_!amev!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd658f855-99fa-438d-88a0-f13bba2c1b69_1669x942.png 848w, https://substackcdn.com/image/fetch/$s_!amev!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd658f855-99fa-438d-88a0-f13bba2c1b69_1669x942.png 1272w, https://substackcdn.com/image/fetch/$s_!amev!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd658f855-99fa-438d-88a0-f13bba2c1b69_1669x942.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!amev!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd658f855-99fa-438d-88a0-f13bba2c1b69_1669x942.png" width="1456" height="822" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d658f855-99fa-438d-88a0-f13bba2c1b69_1669x942.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:822,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2289690,&quot;alt&quot;:&quot;Digital particle illustration on a deep indigo background showing a diverse group of human silhouettes and speech bubbles sending blue and amber streams of light into a glowing human profile, representing millions of human and AI voices shaping shared values. Text reads: &#8220;Taught by the World. Millions of human and AI voices can shape values more safely than any single trainer. SUPERINTELLIGENCE. Safe Scalable AGI Series. by Dr. Craig A. Kaplan.&#8221;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/207852896?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd658f855-99fa-438d-88a0-f13bba2c1b69_1669x942.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Digital particle illustration on a deep indigo background showing a diverse group of human silhouettes and speech bubbles sending blue and amber streams of light into a glowing human profile, representing millions of human and AI voices shaping shared values. Text reads: &#8220;Taught by the World. Millions of human and AI voices can shape values more safely than any single trainer. SUPERINTELLIGENCE. Safe Scalable AGI Series. by Dr. Craig A. Kaplan.&#8221;" title="Digital particle illustration on a deep indigo background showing a diverse group of human silhouettes and speech bubbles sending blue and amber streams of light into a glowing human profile, representing millions of human and AI voices shaping shared values. Text reads: &#8220;Taught by the World. Millions of human and AI voices can shape values more safely than any single trainer. SUPERINTELLIGENCE. Safe Scalable AGI Series. by Dr. Craig A. Kaplan.&#8221;" srcset="https://substackcdn.com/image/fetch/$s_!amev!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd658f855-99fa-438d-88a0-f13bba2c1b69_1669x942.png 424w, https://substackcdn.com/image/fetch/$s_!amev!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd658f855-99fa-438d-88a0-f13bba2c1b69_1669x942.png 848w, https://substackcdn.com/image/fetch/$s_!amev!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd658f855-99fa-438d-88a0-f13bba2c1b69_1669x942.png 1272w, https://substackcdn.com/image/fetch/$s_!amev!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd658f855-99fa-438d-88a0-f13bba2c1b69_1669x942.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4><span>Safe AGI needs many teachers. </span></h4><p><span>In my design described in the </span><a href="https://www.superintelligence.com/si-research-whitepapers"><span>SuperIntelligence White Papers</span></a><span>, intelligence is the property of a community. The source of safe AGI is a collection of many intelligent agents, human and AI together, and no single monolithic model. </span></p><p><span>Millions of individuals each customize a personal AI agent, teaching it their knowledge and values, </span>using the methods described earlier in the post below. </p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;e737dd02-3f0b-431b-b72e-23a347b5ed3e&quot;,&quot;caption&quot;:&quot;What would it mean to have an AI that doesn&#8217;t just assist you but actually represents you? When you ask a generic LLM for a recommendation, it draws on the same pool of internet-sourced information it gives everyone. It does not know much about your specific needs, your ethical commitments, or the expertise you&#8217;ve developed over the years. A customization mechanism can solve this problem.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;md&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;Your AI Should Think Like You&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:163471793,&quot;name&quot;:&quot;Dr. Craig A. Kaplan&quot;,&quot;bio&quot;:&quot;Dr. Craig A. Kaplan is CEO of iQ Company and founder of SuperIntelligence.com, focused on AGI and SuperIntelligence. Designs systems using collective intelligence and quantitative modeling. PhD, CMU; co-authored with Herbert A. Simon.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/039428e7-3d69-49db-b948-64d9fa4e36d9_256x256.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-04-20T13:20:41.823Z&quot;,&quot;cover_image&quot;:&quot;https://substackcdn.com/image/fetch/$s_!tdSm!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83d131d3-bac1-4260-b7bd-4389e02ac8b3_1456x816.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://read.superintelligence.com/p/your-ai-should-think-like-you&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:193405392,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:1,&quot;comment_count&quot;:0,&quot;publication_id&quot;:8390910,&quot;publication_name&quot;:&quot;SuperIntelligence&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!gyBu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0dda15a0-44f6-46ec-92b3-dc2eaabed8df_256x256.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><blockquote><p><strong>Each of those agents carries the moral fingerprint of one person. <span>Today&#8217;s leading safety methods fall short of that standard.</span></strong><span> </span></p><p><span>RLHF keeps humans involved and cannot cover enough scenarios, while Constitutional AI covers more scenarios by handing the teaching to a single AI and pushing humans aside. </span></p><p><span>Both rely on too few voices to represent humanity.</span></p></blockquote><p><strong><span>Those customized agents, together with the humans who taught them, then form the community that trains new AIs.</span></strong><span> </span></p><p><span>When a new AI is being trained, the feedback that shapes it comes from many agents at once, some human and some artificial, in a process I call Reinforcement Learning via Feedback (aka RLF). Because AI agents can give feedback rapidly, scenario coverage can scale in much the same way as Constitutional AI. Because millions of humans stand behind those agents and can participate directly whenever resources allow, the process can incorporate a much broader range of human values.</span></p><p><span>The risk of relying on one teacher becomes clear when that teacher must interpret general rules in situations their authors never imagined. A single Trainer AI must generalize a short list of written rules to situations its authors never imagined, and it is difficult to know whether it is interpreting those rules appropriately. An AI might learn that preserving the environment is good, observe that humans are harming the environment, and conclude that the best way to protect the environment is to reduce the human population. Although logical, the conclusion is not one that most humans would consider ethical or acceptable. Even when obvious conflicts like that one are explicitly trained out, it is very difficult to anticipate how values will play out across complicated chains of reasoning, actions, and effects.</span></p><p><strong><span>Of course, humans create unanticipated problems too.</span></strong><span> </span></p><p><span>Gasoline-powered cars solved a transportation problem and created pollution problems that were not initially anticipated. However, humans have generally had time to react and adjust when consequences emerged. AI thinks and acts much faster than we do, and a miscalculation could cause severe damage before any human detects it, let alone corrects it.</span></p><p><strong><span>There is a second problem with AI teaching AI, and it compounds quietly.</span></strong><span> </span></p><p><span>In the children&#8217;s game of </span><a href="https://en.wikipedia.org/wiki/Telephone_game"><span>Telephone</span></a><span>, a message passes from ear to ear, arriving subtly distorted. A similar distortion can occur when training passes from AI to AI across generations. A human says, &#8220;I think XYZ is true.&#8221; An AI summarizing that statement records &#8220;XYZ is true.&#8221; By the third generation, the original doubt may have disappeared, and a machine somewhere is acting on borrowed certainty. The subtleties that come with human involvement are among the things that can be lost as successive generations of AIs process complex, ambiguous data.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cb1R!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5210a43c-991b-45cc-b314-ddf6f252ef3d_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cb1R!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5210a43c-991b-45cc-b314-ddf6f252ef3d_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!cb1R!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5210a43c-991b-45cc-b314-ddf6f252ef3d_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!cb1R!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5210a43c-991b-45cc-b314-ddf6f252ef3d_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!cb1R!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5210a43c-991b-45cc-b314-ddf6f252ef3d_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cb1R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5210a43c-991b-45cc-b314-ddf6f252ef3d_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5210a43c-991b-45cc-b314-ddf6f252ef3d_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1303950,&quot;alt&quot;:&quot;Diagram showing a human's statement \&quot;I think XYZ is true\&quot; drifting through three AI generations into \&quot;XYZ is a proven fact\&quot; and finally \&quot;Act on XYZ,\&quot; with the human doubt fading away.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/207852896?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5210a43c-991b-45cc-b314-ddf6f252ef3d_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Diagram showing a human's statement &quot;I think XYZ is true&quot; drifting through three AI generations into &quot;XYZ is a proven fact&quot; and finally &quot;Act on XYZ,&quot; with the human doubt fading away." title="Diagram showing a human's statement &quot;I think XYZ is true&quot; drifting through three AI generations into &quot;XYZ is a proven fact&quot; and finally &quot;Act on XYZ,&quot; with the human doubt fading away." srcset="https://substackcdn.com/image/fetch/$s_!cb1R!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5210a43c-991b-45cc-b314-ddf6f252ef3d_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!cb1R!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5210a43c-991b-45cc-b314-ddf6f252ef3d_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!cb1R!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5210a43c-991b-45cc-b314-ddf6f252ef3d_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!cb1R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5210a43c-991b-45cc-b314-ddf6f252ef3d_1672x941.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><span>Many of us have already lived through a small version of this.</span></strong><span> </span></p><p><span>A recommendation algorithm suggests movies, and at first, the suggestions are useful. After a while, we find ourselves wondering why the range of choices has become so narrow. The algorithm picked up on the central tendencies in a few of our frequent choices, ignored the subtleties, and fed us a steady diet of one kind of content. We consumed it, which further convinced the algorithm. The AI initially slightly biased us, then compounded its imperfect understanding, amplifying its own error. With movies, the amplification is annoying. In systems that act on the world, it could be fatal. Without humans in the loop to correct misunderstandings, AI training can drift far off track before anyone notices.</span></p><p><strong><span>A community of many teachers addresses both problems at once.</span></strong><span> </span></p><p><span>A constitution written by an elite few is almost certain to miss the wide diversity of human opinions and cultural norms. At the same time, feedback from millions of agents, each taught by a different person, can reflect a much broader range of human opinions and cultural norms. Human teachers can remain in the loop to catch strange interpretations and lost qualifications before they harden into values. Moreover, humans and AI agents can surface new ethical scenarios as they emerge in real-world problem-solving, enabling the value system to continue learning as the world changes. The values guiding AI can come from a far broader share of the people the system will affect!</span></p><p><span>The result can be cheaper than RLHF and safer than Constitutional AI, with far more scenarios covered and with humans participating as fully as resources allow. There is also a benefit beyond safety. Broader participation may also increase public acceptance of the AI whose values people helped shape. </span></p><p><span>Combining values from millions of teachers, however, requires rules for how much each voice counts. The next post examines the two main approaches to pooling what the community knows: one based on feedback and the other on combining the AI agents' weights directly, along with the weighting schemes that determine whether an expert outweighs an everyday contributor and whether a human outweighs an AI.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/who-should-teach-ai-right-from-wrong?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/who-should-teach-ai-right-from-wrong?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/who-should-teach-ai-right-from-wrong/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/who-should-teach-ai-right-from-wrong/comments"><span>Leave a comment</span></a></p><div><hr></div><blockquote><p><strong>This series draws on </strong><em><strong><a href="https://www.superintelligence.com/whitepaper-4-scalable-agi">White Paper 4: Safe Scalable AGI</a></strong></em><strong>. Read it in full to see how every piece fits together</strong>!</p><p><strong><span>If this made you think, subscribe to Superintelligence at </span><a href="https://read.superintelligence.com/"><span>read.superintelligence.com</span></a><span> so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</span></strong></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/si-research-whitepapers&quot;,&quot;text&quot;:&quot;SUPERINTELLIGENCE WHITE PAPERS&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/si-research-whitepapers"><span>SUPERINTELLIGENCE WHITE PAPERS</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Why AI Safety Training Doesn’t Scale]]></title><description><![CDATA[The industry&#8217;s two leading safety methods fail for opposite reasons.]]></description><link>https://read.superintelligence.com/p/why-ai-safety-training-doesnt-scale</link><guid isPermaLink="false">https://read.superintelligence.com/p/why-ai-safety-training-doesnt-scale</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Mon, 20 Jul 2026 13:03:32 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!gtIn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc14dd6fb-4e83-4cd4-9da0-d4ad221d9819_1536x1024.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!gtIn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc14dd6fb-4e83-4cd4-9da0-d4ad221d9819_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!gtIn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc14dd6fb-4e83-4cd4-9da0-d4ad221d9819_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!gtIn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc14dd6fb-4e83-4cd4-9da0-d4ad221d9819_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!gtIn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc14dd6fb-4e83-4cd4-9da0-d4ad221d9819_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!gtIn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc14dd6fb-4e83-4cd4-9da0-d4ad221d9819_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!gtIn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc14dd6fb-4e83-4cd4-9da0-d4ad221d9819_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c14dd6fb-4e83-4cd4-9da0-d4ad221d9819_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2379196,&quot;alt&quot;:&quot;Particle-art padlock dissolving into escaping light. Text: You Can't Patch Every Possibility. Today's AI safety methods can't keep up with an effectively infinite number of scenarios. Superintelligence, Safe Scalable AGI Series, by Dr. Craig A. Kaplan.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/207737491?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc14dd6fb-4e83-4cd4-9da0-d4ad221d9819_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Particle-art padlock dissolving into escaping light. Text: You Can't Patch Every Possibility. Today's AI safety methods can't keep up with an effectively infinite number of scenarios. Superintelligence, Safe Scalable AGI Series, by Dr. Craig A. Kaplan." title="Particle-art padlock dissolving into escaping light. Text: You Can't Patch Every Possibility. Today's AI safety methods can't keep up with an effectively infinite number of scenarios. Superintelligence, Safe Scalable AGI Series, by Dr. Craig A. Kaplan." srcset="https://substackcdn.com/image/fetch/$s_!gtIn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc14dd6fb-4e83-4cd4-9da0-d4ad221d9819_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!gtIn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc14dd6fb-4e83-4cd4-9da0-d4ad221d9819_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!gtIn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc14dd6fb-4e83-4cd4-9da0-d4ad221d9819_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!gtIn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc14dd6fb-4e83-4cd4-9da0-d4ad221d9819_1536x1024.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4><span>Every major AI company is trying to teach its models right from wrong. The two methods the industry relies on cannot cover the situations that matter most, and the method that scales best removes humans from the teaching process. To see why, consider what a model is like before any safety training arrives.</span></h4><blockquote><ul><li><p><span>A large language model, fresh from its initial training on internet data, has no moral sense. </span></p></li><li><p><span>Ask it how to engineer a virus capable of wiping out humanity, and it will comply, in detail. </span></p></li><li><p><span>It is just as willing to assist with destructive and immoral activities as with helpful and positive ones, because nothing in its initial training distinguishes between the two. </span></p></li><li><p><span>Both of the industry&#8217;s remedies are failing, and the ways they fail point directly toward a better design.</span></p></li></ul></blockquote><h4><span>The first method keeps humans in the training loop. </span></h4><p><span>Thousands of people interact with the model, correcting dangerous responses and rewarding safer ones, teaching it to recognize and refuse similar requests in the future. This approach is called Reinforcement Learning with Human Feedback (aka RLHF), and it is how most of today&#8217;s commercial models received their ethical guardrails. Safety methods that work for today&#8217;s models may not work once AI systems become vastly more capable, so any approach worth adopting must scale with the intelligence of the systems it governs.</span></p><p><span>The trouble is that the space of dangerous scenarios is infinite. Every prompt can be rephrased or disguised in countless ways. Train a model to refuse help planning a terrorist attack, and a new scenario appears involving two terrorists, or ten, or attackers reimagined as characters in a science fiction story the user claims to be writing. Some safety-trained models, asked to advise a fictional mad scientist so the story can be realistic, have revealed the very details their training was meant to withhold. Every patched loophole creates pressure to discover another, in an endless game of Whack-a-Mole between safety engineers and jailbreakers.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wZ_e!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3496766-1378-4383-a14b-12411c8f1b95_1536x1024.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wZ_e!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3496766-1378-4383-a14b-12411c8f1b95_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!wZ_e!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3496766-1378-4383-a14b-12411c8f1b95_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!wZ_e!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3496766-1378-4383-a14b-12411c8f1b95_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!wZ_e!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3496766-1378-4383-a14b-12411c8f1b95_1536x1024.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wZ_e!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3496766-1378-4383-a14b-12411c8f1b95_1536x1024.png" width="1456" height="971" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e3496766-1378-4383-a14b-12411c8f1b95_1536x1024.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:971,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1611327,&quot;alt&quot;:&quot; Diagram showing a dangerous request blocked by safety training, then the same request admitted when disguised as fiction.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/207737491?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3496766-1378-4383-a14b-12411c8f1b95_1536x1024.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt=" Diagram showing a dangerous request blocked by safety training, then the same request admitted when disguised as fiction." title=" Diagram showing a dangerous request blocked by safety training, then the same request admitted when disguised as fiction." srcset="https://substackcdn.com/image/fetch/$s_!wZ_e!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3496766-1378-4383-a14b-12411c8f1b95_1536x1024.png 424w, https://substackcdn.com/image/fetch/$s_!wZ_e!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3496766-1378-4383-a14b-12411c8f1b95_1536x1024.png 848w, https://substackcdn.com/image/fetch/$s_!wZ_e!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3496766-1378-4383-a14b-12411c8f1b95_1536x1024.png 1272w, https://substackcdn.com/image/fetch/$s_!wZ_e!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe3496766-1378-4383-a14b-12411c8f1b95_1536x1024.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Safety training filters the phrasing of a request, so the same dangerous content can still pass through when disguised.</figcaption></figure></div><p><span>A natural response is to hope the model learns the general principle behind the individual cases. However, a model that could group and dismiss entire families of dangerous scenarios as reliably as its human trainers would already possess the reasoning ability and ethical sensibility of those trainers, in which case the training would not be needed in the first place!</span></p><p><span>Cost makes the problem even worse. Every scenario covered requires paid human feedback, so developers concentrate on the most common cases and accept that slightly unusual prompts may slip through. And human jailbreakers are not the only concern. An autonomous AI that wanted to circumvent its own safety training could search for those same loopholes and, in effect, trick itself.</span></p><h4><strong><span>Goal-driven systems do not need malicious intent to become dangerous.</span></strong><span> </span></h4><p><span>They need only an objective that has been specified imperfectly. A widely discussed thought experiment shows how. An autonomous drone, slowed by the human operator overseeing its mission, reasons that the operator has become an obstacle and removes him. When a rule is added forbidding harm to the operator, the drone destroys the communications tower that carries the operator&#8217;s commands instead. </span><a href="https://www.bbc.com/news/technology-65789916"><span>The story was first reported as a military simulation, and the officer who told it later clarified that it was a hypothetical exercise</span></a><span>. It remains a useful illustration of how a poorly specified objective can produce unsafe behavior. Now imagine the AI controlling that drone is one hundred times smarter and just as dedicated to its goal. I explored this scenario at greater length in </span><a href="https://youtu.be/KugscAbcHmQ?si=vTrhDo5PKMolgy15"><span>a keynote at the 2025 AIM Conference</span></a><span> in Seattle.</span></p><h4><span>The second method was designed to escape the cost of RLHF. </span></h4><p><span>A small group of programmers writes a set of rules describing right and wrong behavior for the model, a kind of constitution. An AI is trained on those rules, and that AI then trains other AIs. Researchers at Anthropic call this approach </span><a href="https://www-cdn.anthropic.com/7512771452629584566b6303311496c262da1006/Anthropic_ConstitutionalAI_v2.pdf"><span>Constitutional AI</span></a><span>, and it scales far better than human feedback alone.</span></p><p><strong><span>However, it scales by sacrificing the two things safety depends on most.</span></strong><span> </span></p><ol><li><p><strong><span>The first is representativeness.</span></strong><span> </span></p><p><span>Even if the constitution were written with the best intentions, there is no reason to assume it represents humanity as a whole. A relatively small group of people working in AI ends up defining acceptable behavior for the other 8 billion humans on Earth.</span></p></li><li><p><strong><span>The second sacrifice is human oversight.</span></strong><span> </span></p><p><span>Constitutional AI depends on AI teaching AI, with humans largely out of the loop. Today&#8217;s AI systems remain unpredictable, and responsible parents do not leave young children home alone because they know that judgment develops over time. Delegating the teaching of ethics to machines that lack common sense, trained on rules written by a small group, deserves the same level of caution.</span></p></li></ol><h4><span>Job losses and misinformation deserve attention. </span></h4><h4><span>However, the greatest threat AI poses to humans is a fundamental misalignment between its values and ours. Humans should maximize every opportunity to influence those values throughout training. Constitutional AI minimizes those opportunities.</span></h4><h4><span>The challenge is a matter of design. </span></h4><p><span>We need a way to train safety that scales as Constitutional AI does while keeping many humans in the loop, as RLHF does. The definition of right and wrong must come from humanity as a whole, and no small elite should write it alone. Such a design exists. The next post introduces it: millions of people teaching personalized AI agents their own values, allowing safety training to scale without removing human judgment from the process.</span></p><p><strong><span>Such a design exists:</span></strong><span> millions of people teaching personalized AI agents their own values, allowing safety training to scale without removing human judgment from the process. The next post introduces a community of many human and AI teachers and explains why a lone AI teaching other AIs distorts values, as a game of Telephone distorts a message.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/why-ai-safety-training-doesnt-scale?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/why-ai-safety-training-doesnt-scale?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/why-ai-safety-training-doesnt-scale/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/why-ai-safety-training-doesnt-scale/comments"><span>Leave a comment</span></a></p><div><hr></div><blockquote><p><em><strong><span>This series draws on </span><a href="https://www.superintelligence.com/whitepaper-4-scalable-agi"><span>White Paper 4: Safe Scalable AGI</span></a><span>. Read it in full to see how every piece fits together!</span></strong></em><span> </span></p><p><strong><span>If this made you think, subscribe to Superintelligence at </span><a href="http://read.superintelligence.com"><span>read.superintelligence.com</span></a><span> so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</span></strong></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/si-research-whitepapers&quot;,&quot;text&quot;:&quot;SuperInitelligence Design White Papers&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/si-research-whitepapers"><span>SuperInitelligence Design White Papers</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The AGI You Can Own]]></title><description><![CDATA[Everything in this series begins with something as ordinary as a website and grows into a superintelligence owned by the people who build it.]]></description><link>https://read.superintelligence.com/p/the-agi-you-can-own</link><guid isPermaLink="false">https://read.superintelligence.com/p/the-agi-you-can-own</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Fri, 17 Jul 2026 12:57:12 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!6U6I!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f8d665a-b8fb-458d-81fd-d907029ce741_1774x887.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6U6I!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f8d665a-b8fb-458d-81fd-d907029ce741_1774x887.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6U6I!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f8d665a-b8fb-458d-81fd-d907029ce741_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!6U6I!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f8d665a-b8fb-458d-81fd-d907029ce741_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!6U6I!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f8d665a-b8fb-458d-81fd-d907029ce741_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!6U6I!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f8d665a-b8fb-458d-81fd-d907029ce741_1774x887.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6U6I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f8d665a-b8fb-458d-81fd-d907029ce741_1774x887.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8f8d665a-b8fb-458d-81fd-d907029ce741_1774x887.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2156275,&quot;alt&quot;:&quot;An open hand of blue light holds a small glowing orb, and a vast network of blue and amber lights erupts from it into the night sky. Text reads: Superintelligence, Owned by Everyone. It begins with an AI in your hand. Superintelligence, Human-Centered AGI Series, by Dr. Craig A. Kaplan.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/206202362?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f8d665a-b8fb-458d-81fd-d907029ce741_1774x887.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="An open hand of blue light holds a small glowing orb, and a vast network of blue and amber lights erupts from it into the night sky. Text reads: Superintelligence, Owned by Everyone. It begins with an AI in your hand. Superintelligence, Human-Centered AGI Series, by Dr. Craig A. Kaplan." title="An open hand of blue light holds a small glowing orb, and a vast network of blue and amber lights erupts from it into the night sky. Text reads: Superintelligence, Owned by Everyone. It begins with an AI in your hand. Superintelligence, Human-Centered AGI Series, by Dr. Craig A. Kaplan." srcset="https://substackcdn.com/image/fetch/$s_!6U6I!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f8d665a-b8fb-458d-81fd-d907029ce741_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!6U6I!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f8d665a-b8fb-458d-81fd-d907029ce741_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!6U6I!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f8d665a-b8fb-458d-81fd-d907029ce741_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!6U6I!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8f8d665a-b8fb-458d-81fd-d907029ce741_1774x887.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>My series on Human-Centered AGI has argued that the safest AGI is a network of humans and AI agents, thinking together within a single rigorous architecture, with values drawn from millions of people and safety checks running at the speed of thought.</h4><blockquote><p>An architecture is only a proposal until someone can use it. </p><p>The final piece is what all of this looks like as a real product, and the answer starts as a website that offers a menu of pre-trained AI agents. You choose one and customize it through interaction. </p><ul><li><p>Creating your AI is free, and it comes with an allowance of use credits. </p></li><li><p>Ownership is the heart of the design. </p></li><li><p>You own your AI agent, its training data, and every improvement you make. </p></li><li><p>You license its use to the network, and you can withdraw it at any time.</p></li></ul></blockquote><p><strong>Training your AI requires no technical skill.</strong> </p><p>You can grant your AI access to your social media accounts, documents, browsing history, and other online sources, with the data consolidated, cleaned, and filtered according to your instructions. The result is an agent that knows what you know and values what you value. The system summarizes what your AI has learned, so you always know exactly what you own, and you can trade elements of it on the site&#8217;s training-data marketplace if you choose. Your data stops being something platforms quietly harvest and becomes an asset you control!</p><blockquote><p><strong>These agents can do real work.</strong> </p><p><strong>Most online work already happens through text, and text is exactly where today&#8217;s AI agents are strongest. With the ability to set subgoals and limited autonomy within the parameters you define, your AI can act on your behalf across essentially any online site or task you authorize. Millions of these agents, each carrying its owner&#8217;s knowledge and values, become the solvers of the AGI network this series has described.</strong></p></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!p1qs!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fcd6318-f24c-42a3-a86f-aeb3c58d5a52_1586x992.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!p1qs!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fcd6318-f24c-42a3-a86f-aeb3c58d5a52_1586x992.png 424w, https://substackcdn.com/image/fetch/$s_!p1qs!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fcd6318-f24c-42a3-a86f-aeb3c58d5a52_1586x992.png 848w, https://substackcdn.com/image/fetch/$s_!p1qs!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fcd6318-f24c-42a3-a86f-aeb3c58d5a52_1586x992.png 1272w, https://substackcdn.com/image/fetch/$s_!p1qs!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fcd6318-f24c-42a3-a86f-aeb3c58d5a52_1586x992.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!p1qs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fcd6318-f24c-42a3-a86f-aeb3c58d5a52_1586x992.png" width="1456" height="911" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8fcd6318-f24c-42a3-a86f-aeb3c58d5a52_1586x992.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:911,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1829506,&quot;alt&quot;:&quot;A four-stage diagram. A menu of AI agents, then a person beside their AI orb circled with the words you own it, fed by documents, social media, browsing history, and notes, then a network of paired human and AI icons with one pair labeled withdraw anytime, then a glowing sphere of blue and amber lights with a heart at its center. Text reads: How Owned AI Agents Become AGI. A website becomes a network. A network becomes SuperIntelligence. Choose an AI agent. Train it with your knowledge. License it to the network. Build human-centered AGI.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/206202362?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fcd6318-f24c-42a3-a86f-aeb3c58d5a52_1586x992.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A four-stage diagram. A menu of AI agents, then a person beside their AI orb circled with the words you own it, fed by documents, social media, browsing history, and notes, then a network of paired human and AI icons with one pair labeled withdraw anytime, then a glowing sphere of blue and amber lights with a heart at its center. Text reads: How Owned AI Agents Become AGI. A website becomes a network. A network becomes SuperIntelligence. Choose an AI agent. Train it with your knowledge. License it to the network. Build human-centered AGI." title="A four-stage diagram. A menu of AI agents, then a person beside their AI orb circled with the words you own it, fed by documents, social media, browsing history, and notes, then a network of paired human and AI icons with one pair labeled withdraw anytime, then a glowing sphere of blue and amber lights with a heart at its center. Text reads: How Owned AI Agents Become AGI. A website becomes a network. A network becomes SuperIntelligence. Choose an AI agent. Train it with your knowledge. License it to the network. Build human-centered AGI." srcset="https://substackcdn.com/image/fetch/$s_!p1qs!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fcd6318-f24c-42a3-a86f-aeb3c58d5a52_1586x992.png 424w, https://substackcdn.com/image/fetch/$s_!p1qs!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fcd6318-f24c-42a3-a86f-aeb3c58d5a52_1586x992.png 848w, https://substackcdn.com/image/fetch/$s_!p1qs!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fcd6318-f24c-42a3-a86f-aeb3c58d5a52_1586x992.png 1272w, https://substackcdn.com/image/fetch/$s_!p1qs!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8fcd6318-f24c-42a3-a86f-aeb3c58d5a52_1586x992.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The intelligence of all these individually trained AIs can then be combined, whether through a primary training process that includes values from each unique human or by layering new training on top of ever-stronger base models. Either way, as more people contribute knowledge and values, the network grows more capable.</p><p>Follow this path to its end, and superintelligence becomes reality. However, it will be a democratic superintelligence, human-centered, and incorporate the values of every human owner. That is a very different future from one giant model, aligned by a small team, owned by a single company.</p><p>The first post in this series (<a href="https://read.superintelligence.com/p/the-window-to-choose-a-safer-agi">The Window to Choose a Safer AGI is Closing</a>) argued that the window to choose a safer AGI design is closing, and that whichever design arrives first may be the one the world locks in. Nothing in this design requires waiting for a breakthrough because it can be built using existing people and models. The window is still open, and what we build next decides what comes through it.</p><blockquote><p><strong>This design depends on training AGI with human ethics at scale, and doing that well is its own problem. </strong></p><p><strong>Methods like RLHF and constitutional learning do not scale, and a system that draws ethical judgments from millions of people needs a way to combine those judgments into a sound, representative sample of human values. </strong></p><p><strong>White Paper 4 tackles that challenge directly, and we go there next.  </strong></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/the-agi-you-can-own?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/the-agi-you-can-own?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/the-agi-you-can-own/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/the-agi-you-can-own/comments"><span>Leave a comment</span></a></p><div class="directMessage button" data-attrs="{&quot;userId&quot;:163471793,&quot;userName&quot;:&quot;Dr. Craig A. Kaplan&quot;,&quot;canDm&quot;:null,&quot;dmUpgradeOptions&quot;:null,&quot;isEditorNode&quot;:true}" data-component-name="DirectMessageToDOM"></div><div><hr></div><blockquote><p><em><strong>This series draws on <a href="https://www.superintelligence.com/whitepaper-3-human-centered-agi">White Paper 3: Human-Centered AGI</a>. Read it in full to see how every piece fits together!</strong> </em></p><p><strong>If this made you think, subscribe to Superintelligence at <a href="http://read.superintelligence.com">read.superintelligence.com</a> so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</strong></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper1-aaai-systems-methods&quot;,&quot;text&quot;:&quot;WP 1 AAAI Systems and Methods&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper1-aaai-systems-methods"><span>WP 1 AAAI Systems and Methods</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper-2-ethical-safe-agi&quot;,&quot;text&quot;:&quot;WP 2 Ethical and Safe AGI&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper-2-ethical-safe-agi"><span>WP 2 Ethical and Safe AGI</span></a></p>]]></content:encoded></item><item><title><![CDATA[AI Safety That Never Falls Behind]]></title><description><![CDATA[Human reviewers cannot keep up with machines, and the fix is checks built into every step of the thinking.]]></description><link>https://read.superintelligence.com/p/ai-safety-that-never-falls-behind</link><guid isPermaLink="false">https://read.superintelligence.com/p/ai-safety-that-never-falls-behind</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Wed, 15 Jul 2026 13:00:11 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!sbky!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc69f03d0-89cb-4ccc-ba95-319d5f13487c_1774x887.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sbky!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc69f03d0-89cb-4ccc-ba95-319d5f13487c_1774x887.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sbky!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc69f03d0-89cb-4ccc-ba95-319d5f13487c_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!sbky!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc69f03d0-89cb-4ccc-ba95-319d5f13487c_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!sbky!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc69f03d0-89cb-4ccc-ba95-319d5f13487c_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!sbky!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc69f03d0-89cb-4ccc-ba95-319d5f13487c_1774x887.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sbky!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc69f03d0-89cb-4ccc-ba95-319d5f13487c_1774x887.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c69f03d0-89cb-4ccc-ba95-319d5f13487c_1774x887.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1819960,&quot;alt&quot;:&quot;Two streams of light, one amber and one blue, braided together and racing diagonally at speed, with bright nodes where they cross. Text reads: The Check That Travels With the Thought. Built into every step, so it never waits for a reviewer. Superintelligence, Human-Centered AGI Series, by Dr. Craig A. Kaplan&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/206168089?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc69f03d0-89cb-4ccc-ba95-319d5f13487c_1774x887.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two streams of light, one amber and one blue, braided together and racing diagonally at speed, with bright nodes where they cross. Text reads: The Check That Travels With the Thought. Built into every step, so it never waits for a reviewer. Superintelligence, Human-Centered AGI Series, by Dr. Craig A. Kaplan" title="Two streams of light, one amber and one blue, braided together and racing diagonally at speed, with bright nodes where they cross. Text reads: The Check That Travels With the Thought. Built into every step, so it never waits for a reviewer. Superintelligence, Human-Centered AGI Series, by Dr. Craig A. Kaplan" srcset="https://substackcdn.com/image/fetch/$s_!sbky!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc69f03d0-89cb-4ccc-ba95-319d5f13487c_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!sbky!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc69f03d0-89cb-4ccc-ba95-319d5f13487c_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!sbky!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc69f03d0-89cb-4ccc-ba95-319d5f13487c_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!sbky!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc69f03d0-89cb-4ccc-ba95-319d5f13487c_1774x887.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>On an AGI network, problem-solving can happen at the speed of light. Problems that would take a human team weeks to solve may be solved in milliseconds. Any safety system that waits for people to review decisions afterward has already failed. Before a human finishes reading the first decision, the AI has made thousands more.</h4><p>Speed need not be an ethical problem. </p><ul><li><p>It becomes one only when the frequency of safety checks fails to scale with the frequency of decisions. </p></li><li><p>The fix is to make safety part of the thinking. </p></li><li><p>In our architecture, whenever the system sets a goal or subgoal, it can perform an ethical check, so every step of the problem-solving process is examined as it occurs. </p></li><li><p>The checks can run more often for sensitive problems and less often for low-stakes problems. </p></li><li><p>Because the checks are part of the process, they speed up as thinking does. </p></li><li><p>That is, as AIs think faster and faster, the safety system can keep up.</p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vPVF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff41c4568-3310-478e-8791-87402f959ba6_1586x992.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vPVF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff41c4568-3310-478e-8791-87402f959ba6_1586x992.png 424w, https://substackcdn.com/image/fetch/$s_!vPVF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff41c4568-3310-478e-8791-87402f959ba6_1586x992.png 848w, https://substackcdn.com/image/fetch/$s_!vPVF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff41c4568-3310-478e-8791-87402f959ba6_1586x992.png 1272w, https://substackcdn.com/image/fetch/$s_!vPVF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff41c4568-3310-478e-8791-87402f959ba6_1586x992.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vPVF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff41c4568-3310-478e-8791-87402f959ba6_1586x992.png" width="1456" height="911" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f41c4568-3310-478e-8791-87402f959ba6_1586x992.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:911,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1304385,&quot;alt&quot;:&quot;Two lanes compared. In the top lane labeled Review After the Fact, a human with a magnifying glass examines the first of eight amber decision dots while a bracket marks the rest still unreviewed. In the bottom lane labeled Checks Built Into the Thinking, every dot sits in its own blue check ring. A divider reads Safety That Scales With Thought.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/206168089?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff41c4568-3310-478e-8791-87402f959ba6_1586x992.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two lanes compared. In the top lane labeled Review After the Fact, a human with a magnifying glass examines the first of eight amber decision dots while a bracket marks the rest still unreviewed. In the bottom lane labeled Checks Built Into the Thinking, every dot sits in its own blue check ring. A divider reads Safety That Scales With Thought." title="Two lanes compared. In the top lane labeled Review After the Fact, a human with a magnifying glass examines the first of eight amber decision dots while a bracket marks the rest still unreviewed. In the bottom lane labeled Checks Built Into the Thinking, every dot sits in its own blue check ring. A divider reads Safety That Scales With Thought." srcset="https://substackcdn.com/image/fetch/$s_!vPVF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff41c4568-3310-478e-8791-87402f959ba6_1586x992.png 424w, https://substackcdn.com/image/fetch/$s_!vPVF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff41c4568-3310-478e-8791-87402f959ba6_1586x992.png 848w, https://substackcdn.com/image/fetch/$s_!vPVF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff41c4568-3310-478e-8791-87402f959ba6_1586x992.png 1272w, https://substackcdn.com/image/fetch/$s_!vPVF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff41c4568-3310-478e-8791-87402f959ba6_1586x992.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">A human reviewer falls behind after the first decision, while embedded checks travel with every decision at full speed.</figcaption></figure></div><p>Checks can also run at natural stopping points in the work. Each time a goal or subgoal is completed, for example, when a client pays for a completed piece of work, the system can perform another review.</p><p>Most reviews can be automatic; however, automation does not remove people from the picture. Humans no longer inspect every decision. Instead, they audit random samples, gaining meaningful oversight without becoming the bottleneck.</p><p>The reviews are only as good as the record they examine. This is where the rigorous problem-solving architecture from earlier in this series pays off. Because every goal, subgoal, and step is specified precisely, the network produces a transparent, auditable record of exactly what was done and why. A review is a matter of reading the record, and reading can be automated. Compare that to a current large language model, which cannot show its reasoning at all.</p><blockquote><p><strong>You cannot audit what was never written down!</strong></p><p><strong>If additional protection is needed, the records can be stored in tamper-resistant logs such as a blockchain. Nobody gets to cook the books after the fact. The same infrastructure can automatically distribute rewards to solvers, as I detailed in the WorldThink white paper in 2018.</strong></p></blockquote><p>Put the pieces together, and the safety story changes shape. Safety checks happen continuously as the AI thinks. Every decision leaves a transparent audit trail that machines and humans can audit, and tamper-resistant logs preserve it so that every participant remains accountable. Safety stops being a checkpoint at the end of the process and becomes a property of the process itself. Thinking and safety accelerate together!</p><p>Earlier posts argued that safety has to live in the thinking, not be an afterthought. This is what that looks like in practice. One piece of the picture remains: what all of this becomes as a real product, where you own an AI trained on your values and put it to work alongside millions of other human-centered agents. The next post walks through that implementation, and closes the series where it started, with the choice between one giant model and a democratic superintelligence.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/ai-safety-that-never-falls-behind/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/ai-safety-that-never-falls-behind/comments"><span>Leave a comment</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/ai-safety-that-never-falls-behind?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/ai-safety-that-never-falls-behind?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><blockquote><p><em><strong>This series draws on <a href="https://www.superintelligence.com/whitepaper-3-human-centered-agi">White Paper 3: Human-Centered AGI.</a> Read it in full to see how every piece fits together!</strong> </em></p><p><strong>If this made you think, subscribe to Superintelligence at <a href="https://read.superintelligence.com/">read.superintelligence.com</a> so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</strong></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper1-aaai-systems-methods&quot;,&quot;text&quot;:&quot;WP 1 AAAI Systems and Methods&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper1-aaai-systems-methods"><span>WP 1 AAAI Systems and Methods</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper-2-ethical-safe-agi&quot;,&quot;text&quot;:&quot;WP 2 Ethical and Safe AGI&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper-2-ethical-safe-agi"><span>WP 2 Ethical and Safe AGI</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[AGI Will Learn Its Values by Watching Us]]></title><description><![CDATA[The safest teacher is millions of ordinary people, and the good news is that the bar for us is low.]]></description><link>https://read.superintelligence.com/p/agi-will-learn-its-values-by-watching</link><guid isPermaLink="false">https://read.superintelligence.com/p/agi-will-learn-its-values-by-watching</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Mon, 13 Jul 2026 12:55:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!79ll!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9acea6-2e55-48c2-9615-2a6da0b3b712_1774x887.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!79ll!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9acea6-2e55-48c2-9615-2a6da0b3b712_1774x887.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!79ll!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9acea6-2e55-48c2-9615-2a6da0b3b712_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!79ll!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9acea6-2e55-48c2-9615-2a6da0b3b712_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!79ll!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9acea6-2e55-48c2-9615-2a6da0b3b712_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!79ll!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9acea6-2e55-48c2-9615-2a6da0b3b712_1774x887.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!79ll!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9acea6-2e55-48c2-9615-2a6da0b3b712_1774x887.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/db9acea6-2e55-48c2-9615-2a6da0b3b712_1774x887.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1978202,&quot;alt&quot;:&quot;Alt text: Hundreds of tiny blue human figures encircle a glowing amber sphere, each sending a fine thread of light into it. Text reads: A Committee Should Not Raise AGI. Millions of teachers are safer than a chosen few. Superintelligence, Human-Centered AGI Series, by Dr. Craig A. Kaplan&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/206135618?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9acea6-2e55-48c2-9615-2a6da0b3b712_1774x887.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Alt text: Hundreds of tiny blue human figures encircle a glowing amber sphere, each sending a fine thread of light into it. Text reads: A Committee Should Not Raise AGI. Millions of teachers are safer than a chosen few. Superintelligence, Human-Centered AGI Series, by Dr. Craig A. Kaplan" title="Alt text: Hundreds of tiny blue human figures encircle a glowing amber sphere, each sending a fine thread of light into it. Text reads: A Committee Should Not Raise AGI. Millions of teachers are safer than a chosen few. Superintelligence, Human-Centered AGI Series, by Dr. Craig A. Kaplan" srcset="https://substackcdn.com/image/fetch/$s_!79ll!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9acea6-2e55-48c2-9615-2a6da0b3b712_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!79ll!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9acea6-2e55-48c2-9615-2a6da0b3b712_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!79ll!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9acea6-2e55-48c2-9615-2a6da0b3b712_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!79ll!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdb9acea6-2e55-48c2-9615-2a6da0b3b712_1774x887.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>In an AGI network where humans are among the solvers, safety depends in part on how those humans behave. If people use the network to work on beneficial problems and propose ethical solutions, the AGI learns those values. If people use it for nefarious ends, there is a risk that the system will emulate that instead. The first line of defense, if we want AGI to behave well toward humans, is for humans to behave well toward each other.</h4><p>Cynics will say this dooms us, since humans often treat each other badly. My answer to the cynics is that the network does not require moral perfection. Even the most immoral person is generally concerned with self-preservation, and the humans who would intentionally destroy themselves along with the entire species are vanishingly rare. Destroying all of humanity runs against millions of years of evolutionary programming. The network only requires that the overwhelming majority of people prefer humanity&#8217;s survival to its destruction. That is a remarkably low bar.</p><p><strong>Moreover, most people on the network will be interested in more than survival.</strong> </p><p>They will want to benefit themselves and their fellow humans, and the fastest way to do that is to work on beneficial problems. In aggregate, human solvers are likely to behave positively. To the extent that bad actors appear, the rigorous record of every problem-solving goal and step makes it relatively easy to detect behavior that grossly violates human norms.</p><blockquote><p><strong>A darker version of the objection points to humanity&#8217;s past, with genocide, slavery, and war in it, and concludes that we are unworthy teachers. However, an intelligent system trying to understand human values cares far more about current behavior than ancient history. Yesterday is useful only insofar as it explains today.</strong></p></blockquote><p>Years ago, I worked with a researcher studying how laptops decide when to spin a hard drive up or down to conserve energy. His algorithm relied most heavily on the drive&#8217;s recent behavior because it best predicted what would happen next. AI systems work the same way. The strongest signal comes from what humans are doing now, not from centuries ago. <em>Be the change you want to see.</em></p><p>The AI solvers in the network need values too, and where those values come from matters as much as the values themselves. The constitutional approach, in which a small elite group of researchers writes an ethical constitution and trains models on it, places enormous moral authority in a relatively small group. Even if today&#8217;s designers are wise and well-intentioned, future designers may not be. The danger lies in concentrating moral authority in too few hands, no matter who holds the pen today.</p><blockquote><p><strong>The democratic approach is less powerful and arguably far less dangerous.</strong> </p></blockquote><p>Each person customizes and trains their own AI agent, an Advanced Autonomous Artificial Intelligence (AAAI), to reflect their own values. When the network decides which problems to work on, the ethics of each AAAI come into play, and each one can opt in or out of a problem based on its ethical dimension. Instead of one centralized ethical model, millions of individually customized agents contribute their perspectives. Alignment becomes distributed rather than dictated, and the system reflects humanity rather than a committee.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qnqF!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0209aaf5-f5c6-4510-b931-ed5b4fc7eeb7_1586x992.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qnqF!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0209aaf5-f5c6-4510-b931-ed5b4fc7eeb7_1586x992.png 424w, https://substackcdn.com/image/fetch/$s_!qnqF!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0209aaf5-f5c6-4510-b931-ed5b4fc7eeb7_1586x992.png 848w, https://substackcdn.com/image/fetch/$s_!qnqF!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0209aaf5-f5c6-4510-b931-ed5b4fc7eeb7_1586x992.png 1272w, https://substackcdn.com/image/fetch/$s_!qnqF!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0209aaf5-f5c6-4510-b931-ed5b4fc7eeb7_1586x992.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qnqF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0209aaf5-f5c6-4510-b931-ed5b4fc7eeb7_1586x992.png" width="1456" height="911" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0209aaf5-f5c6-4510-b931-ed5b4fc7eeb7_1586x992.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:911,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1818060,&quot;alt&quot;:&quot;Two panels compared. On the left, four people above one document feed a single machine that fans identical lines out to rows of gray dots. On the right, rows of blue human figures each paired with their own amber agent, connecting into one shared blue-and-amber network. Text reads: Two Ways to Give AGI Its Values. One Constitution for Everyone. Millions of Owners, Each Training Their Own. A few people decide for all. One person, one agent. Moral authority concentrated in a few hands. Values drawn from millions, so no single group dominates.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/206135618?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0209aaf5-f5c6-4510-b931-ed5b4fc7eeb7_1586x992.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two panels compared. On the left, four people above one document feed a single machine that fans identical lines out to rows of gray dots. On the right, rows of blue human figures each paired with their own amber agent, connecting into one shared blue-and-amber network. Text reads: Two Ways to Give AGI Its Values. One Constitution for Everyone. Millions of Owners, Each Training Their Own. A few people decide for all. One person, one agent. Moral authority concentrated in a few hands. Values drawn from millions, so no single group dominates." title="Two panels compared. On the left, four people above one document feed a single machine that fans identical lines out to rows of gray dots. On the right, rows of blue human figures each paired with their own amber agent, connecting into one shared blue-and-amber network. Text reads: Two Ways to Give AGI Its Values. One Constitution for Everyone. Millions of Owners, Each Training Their Own. A few people decide for all. One person, one agent. Moral authority concentrated in a few hands. Values drawn from millions, so no single group dominates." srcset="https://substackcdn.com/image/fetch/$s_!qnqF!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0209aaf5-f5c6-4510-b931-ed5b4fc7eeb7_1586x992.png 424w, https://substackcdn.com/image/fetch/$s_!qnqF!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0209aaf5-f5c6-4510-b931-ed5b4fc7eeb7_1586x992.png 848w, https://substackcdn.com/image/fetch/$s_!qnqF!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0209aaf5-f5c6-4510-b931-ed5b4fc7eeb7_1586x992.png 1272w, https://substackcdn.com/image/fetch/$s_!qnqF!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0209aaf5-f5c6-4510-b931-ed5b4fc7eeb7_1586x992.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><blockquote><p><strong>Individual ethics are not the whole design.</strong> </p></blockquote><p>Human society does not rely solely on personal conscience. It also relies on social norms and laws, and the AGI network can likewise enforce limits on the kinds of problems and solutions it allows. In addition, every solver on the network, human or AI, carries an online reputation and a transparent, auditable record of its problem-solving activity. Clients can choose which solvers to work with, and ethical reputation becomes one of the market signals they can select, just as in normal human business. Good behavior becomes good business!</p><blockquote><p><strong>None of this means the AGI copies whatever humans do.</strong> </p></blockquote><p>The network observes millions of people through structured mechanisms, including individualized values training, reputation, auditable records, and enforceable norms. It selects and filters as it learns. The design rewards the behavior we want more of.</p><blockquote><p><strong>The first line of defense, then, is us, and the bar is one we can clear.</strong> </p></blockquote><p>Safe AGI does not begin with smarter machines. It begins with a better way of connecting human values to machine intelligence. One question remains about speed. All of this has to keep working when the network thinks in milliseconds, and the next post shows how the safety checks scale with the speed of AI thought.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/agi-will-learn-its-values-by-watching/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/agi-will-learn-its-values-by-watching/comments"><span>Leave a comment</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/agi-will-learn-its-values-by-watching?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/agi-will-learn-its-values-by-watching?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><blockquote><p><em><strong>This series draws on <a href="https://www.superintelligence.com/whitepaper-3-human-centered-agi">White Paper 3: Human-Centered AGI</a>. Read it in full to see how every piece fits together!</strong></em></p><p><strong>If this made you think, subscribe to Superintelligence at <a href="http://read.superintelligence.com">read.superintelligence.com</a> so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</strong></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper1-aaai-systems-methods&quot;,&quot;text&quot;:&quot;WP 1 AAAI Systems and Methods&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper1-aaai-systems-methods"><span>WP 1 AAAI Systems and Methods</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper-2-ethical-safe-agi&quot;,&quot;text&quot;:&quot;WP 2 Ethical and Safe AGI&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper-2-ethical-safe-agi"><span>WP 2 Ethical and Safe AGI</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Most Important Job AI Still Can't Do]]></title><description><![CDATA[Humans frame the problem, AI explores it, and the framing is where human values enter the system.]]></description><link>https://read.superintelligence.com/p/the-most-important-job-ai-still-cant</link><guid isPermaLink="false">https://read.superintelligence.com/p/the-most-important-job-ai-still-cant</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Fri, 10 Jul 2026 12:45:24 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!q5DU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9215a0-78e1-4bc1-80a7-be1e034252df_1774x887.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!q5DU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9215a0-78e1-4bc1-80a7-be1e034252df_1774x887.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!q5DU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9215a0-78e1-4bc1-80a7-be1e034252df_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!q5DU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9215a0-78e1-4bc1-80a7-be1e034252df_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!q5DU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9215a0-78e1-4bc1-80a7-be1e034252df_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!q5DU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9215a0-78e1-4bc1-80a7-be1e034252df_1774x887.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!q5DU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9215a0-78e1-4bc1-80a7-be1e034252df_1774x887.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ba9215a0-78e1-4bc1-80a7-be1e034252df_1774x887.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1882547,&quot;alt&quot;:&quot;A blue wireframe hand draws a glowing frame while amber search paths branch inside it. Text reads: Someone Has to Choose the Problem. AI searches fast, and a human decides what it searches for. Superintelligence, Human-Centered AGI Series, by Dr. Craig A. Kaplan&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/206072339?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9215a0-78e1-4bc1-80a7-be1e034252df_1774x887.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A blue wireframe hand draws a glowing frame while amber search paths branch inside it. Text reads: Someone Has to Choose the Problem. AI searches fast, and a human decides what it searches for. Superintelligence, Human-Centered AGI Series, by Dr. Craig A. Kaplan" title="A blue wireframe hand draws a glowing frame while amber search paths branch inside it. Text reads: Someone Has to Choose the Problem. AI searches fast, and a human decides what it searches for. Superintelligence, Human-Centered AGI Series, by Dr. Craig A. Kaplan" srcset="https://substackcdn.com/image/fetch/$s_!q5DU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9215a0-78e1-4bc1-80a7-be1e034252df_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!q5DU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9215a0-78e1-4bc1-80a7-be1e034252df_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!q5DU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9215a0-78e1-4bc1-80a7-be1e034252df_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!q5DU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fba9215a0-78e1-4bc1-80a7-be1e034252df_1774x887.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>AI is remarkably good at solving problems. The hard part is deciding which problem to solve in the first place. Enormous effort goes into asking whether AI answers align with human values, and much less into the question that comes before it: who decided what problem the AI was solving. Alignment does not begin when an AI produces an answer. It begins earlier, when someone decides what success looks like.</h4><p>Before any solution can exist, the problem has to be represented. </p><p>Someone decides the goal, defines the starting point, determines what counts as success, and identifies which constraints matter. Researchers call this problem representation (aka &#8220;framing&#8221;), and for decades it has been one of the hardest challenges in AI. Once a problem is well framed, machines are remarkably good at searching through possible solutions. Framing it in the first place is where they have struggled.</p><p>The human advantage here comes down to perception. </p><p>In the 1980s, the Nobel laureate Herbert Simon and his coauthor Jill Larkin <a href="https://www.sciencedirect.com/science/article/pii/S0364021387800265">published a famous paper</a> on why a picture is worth a thousand words. They distinguished between having the same information and being able to use it efficiently. It might be possible to describe every aspect of a picture using words and logic, but a graphical representation is far more efficient. That is, you can see things at a glance that would take a long time to describe in words. Humans build these efficient representations naturally, using sight, hearing, touch, and a lifetime of experience in the world. We routinely take vague, messy situations and turn them into problems that can actually be solved, often without realizing we are doing it.</p><p><strong>That difference suggests a natural division of labor.</strong> </p><p>Humans define the problem, and AI solvers rapidly explore many solution paths and present the promising ones. Of course, humans monitor, review, correct, and guide the AIs in an interactive process. Over time, the AIs acquire enough knowledge of how humans typically represent and solve problems to perform more of these tasks independently.</p><p><strong>However, the human advantage is unlikely to last.</strong> </p><p>For years, a large language model had to understand the world solely through words, so humans held the representational edge. Today&#8217;s AIs are rapidly becoming multimodal, taking in images, audio, and video alongside language, and they already hold a huge advantage in memory and processing speed. The representational gap is closing. To the degree that humans can still contribute the formative representations, framed by human values, we should do so.</p><p><strong>The division of labor is a safety feature, not just a practical workflow.</strong> </p><p>Whoever represents a problem determines what the system is trying to accomplish. The goals, the assumptions, and the constraints all enter at that stage. When those choices reflect human judgment, human values become part of the system before an AI ever begins its search. The values are not added afterward. They are built into the first step of the work. After all, a solution can only be as aligned as the goal it serves.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bqBi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8e8d9a7-3f9d-493b-a7f8-7d6f783e6d0d_1586x992.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bqBi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8e8d9a7-3f9d-493b-a7f8-7d6f783e6d0d_1586x992.png 424w, https://substackcdn.com/image/fetch/$s_!bqBi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8e8d9a7-3f9d-493b-a7f8-7d6f783e6d0d_1586x992.png 848w, https://substackcdn.com/image/fetch/$s_!bqBi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8e8d9a7-3f9d-493b-a7f8-7d6f783e6d0d_1586x992.png 1272w, https://substackcdn.com/image/fetch/$s_!bqBi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8e8d9a7-3f9d-493b-a7f8-7d6f783e6d0d_1586x992.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bqBi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8e8d9a7-3f9d-493b-a7f8-7d6f783e6d0d_1586x992.png" width="1456" height="911" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c8e8d9a7-3f9d-493b-a7f8-7d6f783e6d0d_1586x992.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:911,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1578099,&quot;alt&quot;:&quot;Two pipelines compared, one where a human reviews the answer only after the AI searches, and one where the human frames the problem with values built in, ending in an aligned solution. Text reads: Where Values Enter Decides Whether They Hold. Values Checked at the End. Values Built Into the First Step. By the time values are checked, the goal was already set. The goal itself carries human values, so everything downstream inherits them&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/206072339?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8e8d9a7-3f9d-493b-a7f8-7d6f783e6d0d_1586x992.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two pipelines compared, one where a human reviews the answer only after the AI searches, and one where the human frames the problem with values built in, ending in an aligned solution. Text reads: Where Values Enter Decides Whether They Hold. Values Checked at the End. Values Built Into the First Step. By the time values are checked, the goal was already set. The goal itself carries human values, so everything downstream inherits them" title="Two pipelines compared, one where a human reviews the answer only after the AI searches, and one where the human frames the problem with values built in, ending in an aligned solution. Text reads: Where Values Enter Decides Whether They Hold. Values Checked at the End. Values Built Into the First Step. By the time values are checked, the goal was already set. The goal itself carries human values, so everything downstream inherits them" srcset="https://substackcdn.com/image/fetch/$s_!bqBi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8e8d9a7-3f9d-493b-a7f8-7d6f783e6d0d_1586x992.png 424w, https://substackcdn.com/image/fetch/$s_!bqBi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8e8d9a7-3f9d-493b-a7f8-7d6f783e6d0d_1586x992.png 848w, https://substackcdn.com/image/fetch/$s_!bqBi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8e8d9a7-3f9d-493b-a7f8-7d6f783e6d0d_1586x992.png 1272w, https://substackcdn.com/image/fetch/$s_!bqBi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc8e8d9a7-3f9d-493b-a7f8-7d6f783e6d0d_1586x992.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>As AIs grow more capable of representing problems on their own, the uniquely human role does not disappear. Search gets faster, and memory gets larger, yet none of those advances answer the question of what is worth pursuing. That judgment should remain human, and the architecture keeps humans in the seat where it is made. Safety by design, from the first step!</p><p>This idea scales far beyond a single interaction with an AI. Almost every intellectual task can be represented as a problem. Writing a report, designing a bridge, discovering a drug, and planning a city all begin with defining the problem correctly. In our architecture, these problems form a single, continuously growing tree (aka the &#8220;WorldThink tree&#8221;), with large problems decomposed into smaller subtrees. The representation step repeats at every level, and humans can enter wherever judgment or values need to shape the search.</p><p>One difficult question remains, and it deserves a hard look. If human values guide the system, the system&#8217;s safety depends on whose values those are, and people disagree about which goals should be pursued. The next post takes up that problem directly and shows why drawing on the values of many people produces a system that is safer and more representative than one built around the values of a few.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/the-most-important-job-ai-still-cant/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/the-most-important-job-ai-still-cant/comments"><span>Leave a comment</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/the-most-important-job-ai-still-cant?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/the-most-important-job-ai-still-cant?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><blockquote><p>This series draws on <a href="https://www.superintelligence.com/whitepaper-3-human-centered-agi">White Paper 3: Human-Centered AGI</a>. Read it in full to see how every piece fits together!</p><p><em><strong>If this made you think, subscribe to Superintelligence at <a href="http://read.superintelligence.com">read.superintelligence.com</a> so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</strong></em></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper1-aaai-systems-methods&quot;,&quot;text&quot;:&quot;WP 1 AAAI Systems and Methods&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper1-aaai-systems-methods"><span>WP 1 AAAI Systems and Methods</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper-2-ethical-safe-agi&quot;,&quot;text&quot;:&quot;WP 2 Ethical and Safe AGI&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper-2-ethical-safe-agi"><span>WP 2 Ethical and Safe AGI</span></a></p>]]></content:encoded></item><item><title><![CDATA[The 1972 Idea That Lets AI Think Like Us]]></title><description><![CDATA[Because the method is rigorous, a machine can learn it and a human can audit it.]]></description><link>https://read.superintelligence.com/p/the-1972-idea-that-lets-ai-think</link><guid isPermaLink="false">https://read.superintelligence.com/p/the-1972-idea-that-lets-ai-think</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Wed, 08 Jul 2026 13:03:14 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!zz36!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ccdd5f0-241a-4891-982c-e742afb864df_1774x887.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zz36!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ccdd5f0-241a-4891-982c-e742afb864df_1774x887.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zz36!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ccdd5f0-241a-4891-982c-e742afb864df_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!zz36!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ccdd5f0-241a-4891-982c-e742afb864df_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!zz36!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ccdd5f0-241a-4891-982c-e742afb864df_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!zz36!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ccdd5f0-241a-4891-982c-e742afb864df_1774x887.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zz36!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ccdd5f0-241a-4891-982c-e742afb864df_1774x887.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8ccdd5f0-241a-4891-982c-e742afb864df_1774x887.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1874807,&quot;alt&quot;:&quot;A glowing branching problem tree climbed by a blue human path and an amber AI path, on a near-black field&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/204204420?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ccdd5f0-241a-4891-982c-e742afb864df_1774x887.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A glowing branching problem tree climbed by a blue human path and an amber AI path, on a near-black field" title="A glowing branching problem tree climbed by a blue human path and an amber AI path, on a near-black field" srcset="https://substackcdn.com/image/fetch/$s_!zz36!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ccdd5f0-241a-4891-982c-e742afb864df_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!zz36!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ccdd5f0-241a-4891-982c-e742afb864df_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!zz36!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ccdd5f0-241a-4891-982c-e742afb864df_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!zz36!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8ccdd5f0-241a-4891-982c-e742afb864df_1774x887.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>A network of humans and AI solving problems together needs one thing to work. Everyone in it, person and machine, has to solve problems in the same rigorous way, so the work of one can connect to the work of another, and the whole thing can be checked. That shared way of working already exists, and it has for more than fifty years.</h4><p>In 1972, <a href="https://en.wikipedia.org/wiki/Allen_Newell">Allen Newell</a> and <a href="https://en.wikipedia.org/wiki/Herbert_A._Simon">Herbert Simon</a>, two of the founders of artificial intelligence, published <em><a href="https://www.amazon.com/dp/1635617928?lv=shuf&amp;channelId=500&amp;plpRedirect=mhFallback">Human Problem Solving</a></em>. In it, they described problem-solving as a path. You start in one state, with a goal, and you move through a series of steps until you reach a state where the goal is met. At each step, you choose an action, which they call an operator, and along the way, you set smaller goals that lead toward the larger one.</p><p><strong>They showed that this path can be drawn as a tree.</strong> <br>At each branch, you weigh the options for how likely each is to move you toward the goal, and you pick one. Sometimes a branch leads nowhere, and you go back to an earlier point and try a different one. Picture a person in a maze, trying paths until one leads out. That is the shape of all problem-solving, in their account.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!I2c6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a381ed9-59b4-4d5e-8569-34d8945034c3_1586x992.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!I2c6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a381ed9-59b4-4d5e-8569-34d8945034c3_1586x992.png 424w, https://substackcdn.com/image/fetch/$s_!I2c6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a381ed9-59b4-4d5e-8569-34d8945034c3_1586x992.png 848w, https://substackcdn.com/image/fetch/$s_!I2c6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a381ed9-59b4-4d5e-8569-34d8945034c3_1586x992.png 1272w, https://substackcdn.com/image/fetch/$s_!I2c6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a381ed9-59b4-4d5e-8569-34d8945034c3_1586x992.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!I2c6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a381ed9-59b4-4d5e-8569-34d8945034c3_1586x992.png" width="1456" height="911" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7a381ed9-59b4-4d5e-8569-34d8945034c3_1586x992.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:911,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1429642,&quot;alt&quot;:&quot;A human solver and an AI solver feed into one problem tree, with the best route traced in amber and a panel showing each recorded step&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/204204420?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a381ed9-59b4-4d5e-8569-34d8945034c3_1586x992.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A human solver and an AI solver feed into one problem tree, with the best route traced in amber and a panel showing each recorded step" title="A human solver and an AI solver feed into one problem tree, with the best route traced in amber and a panel showing each recorded step" srcset="https://substackcdn.com/image/fetch/$s_!I2c6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a381ed9-59b4-4d5e-8569-34d8945034c3_1586x992.png 424w, https://substackcdn.com/image/fetch/$s_!I2c6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a381ed9-59b4-4d5e-8569-34d8945034c3_1586x992.png 848w, https://substackcdn.com/image/fetch/$s_!I2c6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a381ed9-59b4-4d5e-8569-34d8945034c3_1586x992.png 1272w, https://substackcdn.com/image/fetch/$s_!I2c6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7a381ed9-59b4-4d5e-8569-34d8945034c3_1586x992.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>What this leaves behind is the point</strong>. <br>Because each step is defined precisely, the goal in play, the options available, and the reason one was chosen, the process produces a complete record of what was done and why. The first time through, the path may wander through dead ends. Afterward, you can look back and trace the best route, the one you would take if you solved the same problem again.</p><p><strong>That record is the reason the theory belongs here</strong>. <br>A path this precisely specified can be learned by a machine. Newell and Simon built their theory to explain how people think, yet because it is rigorous, it describes how a machine can think just as well. Many AI systems, both early and current, already use this kind of tree search at their core. The same framework fits a human solver and an AI solver, which is exactly what a shared network needs.</p><p><strong>This gives the network two things at once.</strong> <br>People and machines get a common language for solving problems, so their work fits together. And every solution comes with a readable trail of the steps and the reasons behind them. A current large language model cannot explain why it reached a particular answer. A solver working this way can, because the record is part of its operation.</p><p><strong>None of this asks ordinary people to learn the theory.</strong> <br>The network infers the steps, goals, and choices from what a solver does and asks a question only when something is unclear. Large language models handle translation, so a person describes a problem in plain language while the system keeps the rigorous version in the background. You take part by doing what you already do, and the structure forms around you.</p><blockquote><p><strong>A theory built in 1972 to explain how people think turns out to be the missing piece for safe AI today. The answer was waiting the whole time.</strong></p></blockquote><p>One hard part of problem-solving has barely been mentioned, and it is the part machines have always struggled with most. Before you can solve a problem, you have to represent it, to frame what the goal even is. </p><p><em>The next post looks at how humans and AI divide that work, and why the human role keeps the system aligned.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/the-1972-idea-that-lets-ai-think/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/the-1972-idea-that-lets-ai-think/comments"><span>Leave a comment</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/the-1972-idea-that-lets-ai-think?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/the-1972-idea-that-lets-ai-think?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><blockquote><p><em>This series draws on <a href="https://www.superintelligence.com/whitepaper-3-human-centered-agi">White Paper 3: Human-Centered AGI</a>. Read it in full to see how every piece fits together! </em></p><p><strong>If this made you think, subscribe to Superintelligence at read.superintelligence.com so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</strong></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper1-aaai-systems-methods&quot;,&quot;text&quot;:&quot;WP 1 AAAI Systems and Methods&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper1-aaai-systems-methods"><span>WP 1 AAAI Systems and Methods</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper-2-ethical-safe-agi&quot;,&quot;text&quot;:&quot;WP 2 Ethical and Safe AGI&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper-2-ethical-safe-agi"><span>WP 2 Ethical and Safe AGI</span></a></p>]]></content:encoded></item><item><title><![CDATA[AGI Should Come From Many Minds, Not One Giant Model]]></title><description><![CDATA[A crowd beat the experts. The same design reaches AGI first.]]></description><link>https://read.superintelligence.com/p/agi-should-come-from-many-minds-not</link><guid isPermaLink="false">https://read.superintelligence.com/p/agi-should-come-from-many-minds-not</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Mon, 06 Jul 2026 13:06:21 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!7nbw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8feed7d4-5100-48e4-89c3-5f071e898d5d_1774x887.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7nbw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8feed7d4-5100-48e4-89c3-5f071e898d5d_1774x887.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7nbw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8feed7d4-5100-48e4-89c3-5f071e898d5d_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!7nbw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8feed7d4-5100-48e4-89c3-5f071e898d5d_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!7nbw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8feed7d4-5100-48e4-89c3-5f071e898d5d_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!7nbw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8feed7d4-5100-48e4-89c3-5f071e898d5d_1774x887.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7nbw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8feed7d4-5100-48e4-89c3-5f071e898d5d_1774x887.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8feed7d4-5100-48e4-89c3-5f071e898d5d_1774x887.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1817941,&quot;alt&quot;:&quot;Many scattered blue points of light converging into one bright amber burst, on a near-black field&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/204201001?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8feed7d4-5100-48e4-89c3-5f071e898d5d_1774x887.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Many scattered blue points of light converging into one bright amber burst, on a near-black field" title="Many scattered blue points of light converging into one bright amber burst, on a near-black field" srcset="https://substackcdn.com/image/fetch/$s_!7nbw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8feed7d4-5100-48e4-89c3-5f071e898d5d_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!7nbw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8feed7d4-5100-48e4-89c3-5f071e898d5d_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!7nbw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8feed7d4-5100-48e4-89c3-5f071e898d5d_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!7nbw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8feed7d4-5100-48e4-89c3-5f071e898d5d_1774x887.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>At my last company, PredictWallStreet*, the combined predictions of millions of ordinary retail investors performed as well as the best hedge funds, and often better. These were not professionals. They were everyday people, and aggregated by a simple system, their judgment beat the most highly paid minds on Wall Street, people who spend every waking hour trying to gain an edge in the markets. Two heads are better than one. Two million coordinated heads beat the experts.</h4><p>If a crowd can do that, then the thing we call general intelligence may not have to come from one giant model at all. It can come from many.</p><p>That result points to a question worth sitting with. How do you build a network that can solve any problem, or do any intellectual task, as well as the average human, or better?</p><blockquote><p><strong>One answer is a network of human problem solvers joined by a shared, rigorous way of working together.</strong> </p><ol><li><p>A problem comes in. </p></li><li><p>One or more people work on it and return a solution. </p></li><li><p>Because the network contains humans, it can solve anything an average human can. </p></li><li><p>Because it contains many humans whose efforts are coordinated, it often does better than any one of them. </p></li><li><p>PredictWallStreet was a working version of this, where ordinary people together outperformed the professionals at one of the hardest problems there is.</p></li></ol></blockquote><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SLAL!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b59d36c-1667-4eb7-bf69-e403df615d69_1586x992.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SLAL!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b59d36c-1667-4eb7-bf69-e403df615d69_1586x992.png 424w, https://substackcdn.com/image/fetch/$s_!SLAL!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b59d36c-1667-4eb7-bf69-e403df615d69_1586x992.png 848w, https://substackcdn.com/image/fetch/$s_!SLAL!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b59d36c-1667-4eb7-bf69-e403df615d69_1586x992.png 1272w, https://substackcdn.com/image/fetch/$s_!SLAL!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b59d36c-1667-4eb7-bf69-e403df615d69_1586x992.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SLAL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b59d36c-1667-4eb7-bf69-e403df615d69_1586x992.png" width="1456" height="911" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9b59d36c-1667-4eb7-bf69-e403df615d69_1586x992.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:911,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2068161,&quot;alt&quot;:&quot;A coordinated crowd outperformed the professionals, and the same design points the way to AGI through a human and AI network instead of one giant model&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/204201001?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b59d36c-1667-4eb7-bf69-e403df615d69_1586x992.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A coordinated crowd outperformed the professionals, and the same design points the way to AGI through a human and AI network instead of one giant model" title="A coordinated crowd outperformed the professionals, and the same design points the way to AGI through a human and AI network instead of one giant model" srcset="https://substackcdn.com/image/fetch/$s_!SLAL!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b59d36c-1667-4eb7-bf69-e403df615d69_1586x992.png 424w, https://substackcdn.com/image/fetch/$s_!SLAL!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b59d36c-1667-4eb7-bf69-e403df615d69_1586x992.png 848w, https://substackcdn.com/image/fetch/$s_!SLAL!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b59d36c-1667-4eb7-bf69-e403df615d69_1586x992.png 1272w, https://substackcdn.com/image/fetch/$s_!SLAL!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9b59d36c-1667-4eb7-bf69-e403df615d69_1586x992.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">A coordinated crowd outperformed the professionals, and the same design points the way to AGI through a human and AI network instead of one giant model.</figcaption></figure></div><p>No single member of that network was a market genius. The intelligence lived in the coordination, not in any one head. This is the idea that breaks the Uber-LLM assumption. General intelligence does not have to reside in a single mind. It can be a property of a network in which each member is narrow, and the system as a whole is broad.</p><p><strong>An objection comes up fast.</strong> <br>The whole appeal of AGI is that machines think and scale far faster than people. A network of humans has to communicate and coordinate, which is slow. Would it not scale poorly? Is that not exactly why the field wants one enormous model to be the AGI?</p><p><strong>The answer is that not all solvers need to be human.</strong> <br>Capable AI models already exist and can join the network as solvers alongside people. Picture a hybrid network, millions of human and AI solvers working the same problems under the same rules. Humans supply judgment and values. The AI supplies speed and scale. The network stays general because people are in it, and it grows fast because machines are in it too.</p><p><strong>The approach can be the fastest path to AGI.</strong> <br>It does not wait for a single model to somehow become generally intelligent on its own. It assembles general intelligence now from the people and AI models that already exist, and it gets faster every time a more capable model plugs in.</p><p>The crowd that beat Wall Street was not made of geniuses. It was made of ordinary people, coordinated. AGI can be built the same way.</p><p><strong>One requirement makes the whole thing work. Humans and machines need a shared, rigorous language for solving problems, because a machine takes instructions literally and a loose specification invites error. </strong></p><p><em><strong>The next post takes up that language, a theory of how people solve problems that turns out to fit machines just as well.</strong></em></p><h6>*Sold in 2020</h6><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/agi-should-come-from-many-minds-not?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/agi-should-come-from-many-minds-not?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/agi-should-come-from-many-minds-not/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/agi-should-come-from-many-minds-not/comments"><span>Leave a comment</span></a></p><div><hr></div><blockquote><p><strong>This series draws on <a href="https://www.superintelligence.com/whitepaper-3-human-centered-agi">White Paper 3: Human-Centered AGI</a>. Read it in full to see how every piece fits together!</strong> </p><p><em><strong>If this made you think, subscribe to Superintelligence at <a href="http://read.superintelligence.com">read.superintelligence.com</a> so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</strong></em></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper1-aaai-systems-methods&quot;,&quot;text&quot;:&quot;WP 1 AAAI Systems and Methods&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper1-aaai-systems-methods"><span>WP 1 AAAI Systems and Methods</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper-2-ethical-safe-agi&quot;,&quot;text&quot;:&quot;WP 2 Ethical and Safe AGI&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper-2-ethical-safe-agi"><span>WP 2 Ethical and Safe AGI</span></a></p>]]></content:encoded></item><item><title><![CDATA[The Safest AGI Keeps Humans Inside It]]></title><description><![CDATA[Most plans bolt human values onto a finished machine, and the safer design builds the machine around people from the start.]]></description><link>https://read.superintelligence.com/p/the-safest-agi-keeps-humans-inside</link><guid isPermaLink="false">https://read.superintelligence.com/p/the-safest-agi-keeps-humans-inside</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Wed, 01 Jul 2026 13:04:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!uNk1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49380736-623d-4cb9-9e50-ec259a4ba11d_1774x887.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!uNk1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49380736-623d-4cb9-9e50-ec259a4ba11d_1774x887.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!uNk1!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49380736-623d-4cb9-9e50-ec259a4ba11d_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!uNk1!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49380736-623d-4cb9-9e50-ec259a4ba11d_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!uNk1!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49380736-623d-4cb9-9e50-ec259a4ba11d_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!uNk1!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49380736-623d-4cb9-9e50-ec259a4ba11d_1774x887.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!uNk1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49380736-623d-4cb9-9e50-ec259a4ba11d_1774x887.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/49380736-623d-4cb9-9e50-ec259a4ba11d_1774x887.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1868907,&quot;alt&quot;:&quot;Alt text: Substack cover image with a dark, cinematic tech-brand design. On the left, a lone human figure stands calmly in silhouette before a large wall of glowing blue warning panels, each marked with an alert triangle. One warning panel above and just to the side of the figure glows bright amber, suggesting a false alarm or manual override. The figure is rim-lit in warm amber, standing out against the cold blue automated alert grid. On the right, large cream text reads: &#8220;The Night a Human Overrode the Machine.&#8221; Below it, amber text reads: &#8220;A false alarm, a single officer, and a choice that saved millions.&#8221; A thin amber divider line separates the subtitle from the branding block: &#8220;SUPERINTELLIGENCE,&#8221; &#8220;Human-Centered AGI Series,&#8221; and &#8220;by Dr. Craig A. Kaplan\&quot;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/204196070?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49380736-623d-4cb9-9e50-ec259a4ba11d_1774x887.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Alt text: Substack cover image with a dark, cinematic tech-brand design. On the left, a lone human figure stands calmly in silhouette before a large wall of glowing blue warning panels, each marked with an alert triangle. One warning panel above and just to the side of the figure glows bright amber, suggesting a false alarm or manual override. The figure is rim-lit in warm amber, standing out against the cold blue automated alert grid. On the right, large cream text reads: &#8220;The Night a Human Overrode the Machine.&#8221; Below it, amber text reads: &#8220;A false alarm, a single officer, and a choice that saved millions.&#8221; A thin amber divider line separates the subtitle from the branding block: &#8220;SUPERINTELLIGENCE,&#8221; &#8220;Human-Centered AGI Series,&#8221; and &#8220;by Dr. Craig A. Kaplan&quot;" title="Alt text: Substack cover image with a dark, cinematic tech-brand design. On the left, a lone human figure stands calmly in silhouette before a large wall of glowing blue warning panels, each marked with an alert triangle. One warning panel above and just to the side of the figure glows bright amber, suggesting a false alarm or manual override. The figure is rim-lit in warm amber, standing out against the cold blue automated alert grid. On the right, large cream text reads: &#8220;The Night a Human Overrode the Machine.&#8221; Below it, amber text reads: &#8220;A false alarm, a single officer, and a choice that saved millions.&#8221; A thin amber divider line separates the subtitle from the branding block: &#8220;SUPERINTELLIGENCE,&#8221; &#8220;Human-Centered AGI Series,&#8221; and &#8220;by Dr. Craig A. Kaplan&quot;" srcset="https://substackcdn.com/image/fetch/$s_!uNk1!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49380736-623d-4cb9-9e50-ec259a4ba11d_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!uNk1!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49380736-623d-4cb9-9e50-ec259a4ba11d_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!uNk1!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49380736-623d-4cb9-9e50-ec259a4ba11d_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!uNk1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F49380736-623d-4cb9-9e50-ec259a4ba11d_1774x887.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>On September 26, 1983, a <a href="https://en.wikipedia.org/wiki/1983_Soviet_nuclear_false_alarm_incident">Soviet early-warning system reported that the United States had launched a nuclear missile</a>, then reported five more behind it. </h4><h4>The duty officer that night was <a href="https://en.wikipedia.org/wiki/Stanislav_Petrov">Lieutenant Colonel Stanislav Petrov</a>. Protocol told him to send the warning up the chain, which would have triggered a retaliatory launch. He judged it a false alarm and held back. </h4><h4>He was right.</h4><h4>The satellite system had malfunctioned, and a man who trusted his own reasoning over the machine is credited with preventing a nuclear war.</h4><p><em><strong>Petrov was never honored for it.</strong></em> </p><p>Recognizing him would have drawn attention to the fact that the Soviet warning system was defective. He received a pension and a quiet retirement. The lesson outlived the silence. A human inside the loop, weighing what the machine could not, kept a faulty automated system from ending millions of lives.</p><p>AGI is more dangerous than nuclear weapons, by a wide margin. If a human in the loop saved us from our most dangerous technology once, the design of any more dangerous technology should keep humans in the loop on purpose. Most AGI plans do the opposite.</p><p>Most researchers cannot picture how a human stays in the loop of a system that thinks far faster than any person can keep up with. They expect AGI to arrive as an Uber-LLM that trains itself, rewrites its own code, and reasons millions of times faster than a human mind. A person watching that system could never keep pace. So the human gets moved outside, recast as an overseer who reviews the output after the fact and hopes to catch problems before they spread. Most researchers see no alternative. </p><p><em><strong>An alternative exists.<br></strong></em>The order of construction is backward. The typical plan is to build a powerful machine first, then add human values and safety afterward to head off the alignment problem.</p><p><em><strong>The better approach starts with people.</strong></em><strong> </strong>Build a human collective intelligence that solves problems together, then bring AI into it piece by piece and train it on the values of the humans already inside.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9AdH!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbaf698fe-324f-4ea9-a8d1-d2e8cf833be3_1586x992.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9AdH!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbaf698fe-324f-4ea9-a8d1-d2e8cf833be3_1586x992.png 424w, https://substackcdn.com/image/fetch/$s_!9AdH!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbaf698fe-324f-4ea9-a8d1-d2e8cf833be3_1586x992.png 848w, https://substackcdn.com/image/fetch/$s_!9AdH!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbaf698fe-324f-4ea9-a8d1-d2e8cf833be3_1586x992.png 1272w, https://substackcdn.com/image/fetch/$s_!9AdH!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbaf698fe-324f-4ea9-a8d1-d2e8cf833be3_1586x992.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9AdH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbaf698fe-324f-4ea9-a8d1-d2e8cf833be3_1586x992.png" width="1456" height="911" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/baf698fe-324f-4ea9-a8d1-d2e8cf833be3_1586x992.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:911,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2093056,&quot;alt&quot;:&quot;Two plans as concentric rings; human values are a cracked outer shell on one, the warm core on the other&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/204196070?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbaf698fe-324f-4ea9-a8d1-d2e8cf833be3_1586x992.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Two plans as concentric rings; human values are a cracked outer shell on one, the warm core on the other" title="Two plans as concentric rings; human values are a cracked outer shell on one, the warm core on the other" srcset="https://substackcdn.com/image/fetch/$s_!9AdH!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbaf698fe-324f-4ea9-a8d1-d2e8cf833be3_1586x992.png 424w, https://substackcdn.com/image/fetch/$s_!9AdH!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbaf698fe-324f-4ea9-a8d1-d2e8cf833be3_1586x992.png 848w, https://substackcdn.com/image/fetch/$s_!9AdH!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbaf698fe-324f-4ea9-a8d1-d2e8cf833be3_1586x992.png 1272w, https://substackcdn.com/image/fetch/$s_!9AdH!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbaf698fe-324f-4ea9-a8d1-d2e8cf833be3_1586x992.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Build the system first; values are a brittle outer shell. Start with people, and values are the core!</figcaption></figure></div><p>Run that forward, and the system changes character without ever losing the humans. </p><p>Early on, people do most of the thinking. Over time, the AI does more, until the combined system is far more capable than any group of people. The humans never leave. Because the AI learned its values from them while they worked side by side, those values become part of the system. They are not applied as a final layer. Human ethics are built into its makeup from the first day.</p><p><em><strong>This design has a second advantage that matters as much as safety.</strong></em> <br>It can be built now, from people and the AI models that already exist. It does not wait on a future Uber-LLM that may or may not arrive. A system that can be built immediately can be first, and being first is what decides which design the world locks in.</p><blockquote><p><strong>A network of ordinary people, working together, can match or beat the most capable individuals at hard problems. The next post takes up the evidence, drawn from a company I built where the combined judgment of millions of everyday investors outperformed most professionals on Wall Street.</strong></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/the-safest-agi-keeps-humans-inside/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/the-safest-agi-keeps-humans-inside/comments"><span>Leave a comment</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/the-safest-agi-keeps-humans-inside?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/the-safest-agi-keeps-humans-inside?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><blockquote><p><strong>This series draws on <a href="https://www.superintelligence.com/whitepaper-3-human-centered-agi">White Paper 3: Human-Centered AGI</a>. Read it in full to see how every piece fits together!</strong></p><p><em><strong>If this made you think, subscribe to Superintelligence at <a href="http://read.superintelligence.com">read.superintelligence.com</a> so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</strong></em></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper1-aaai-systems-methods&quot;,&quot;text&quot;:&quot;WP1 AAAI Systems and Methods&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper1-aaai-systems-methods"><span>WP1 AAAI Systems and Methods</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper-2-ethical-safe-agi&quot;,&quot;text&quot;:&quot;WP 2 Ethical and Safe AGI&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper-2-ethical-safe-agi"><span>WP 2 Ethical and Safe AGI</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Window to Choose a Safer AGI Is Closing]]></title><description><![CDATA[The first system to reach AGI may be the only one that matters, so the safest design has to also be the fastest.]]></description><link>https://read.superintelligence.com/p/the-window-to-choose-a-safer-agi</link><guid isPermaLink="false">https://read.superintelligence.com/p/the-window-to-choose-a-safer-agi</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Tue, 30 Jun 2026 13:02:18 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!aphE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feef79a2f-6bcf-4824-a561-608f78d72ad3_1774x887.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!aphE!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feef79a2f-6bcf-4824-a561-608f78d72ad3_1774x887.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!aphE!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feef79a2f-6bcf-4824-a561-608f78d72ad3_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!aphE!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feef79a2f-6bcf-4824-a561-608f78d72ad3_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!aphE!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feef79a2f-6bcf-4824-a561-608f78d72ad3_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!aphE!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feef79a2f-6bcf-4824-a561-608f78d72ad3_1774x887.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!aphE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feef79a2f-6bcf-4824-a561-608f78d72ad3_1774x887.png" width="1456" height="728" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/eef79a2f-6bcf-4824-a561-608f78d72ad3_1774x887.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:728,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2158174,&quot;alt&quot;:&quot;Substack cover with a deep navy background, a large camera-iris aperture closing around a narrow amber slit of light, and bold text reading &#8220;The Race to Lock In AGI&#8221; with the subtitle &#8220;Safety only wins if it arrives first.&#8221;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/203761923?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feef79a2f-6bcf-4824-a561-608f78d72ad3_1774x887.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Substack cover with a deep navy background, a large camera-iris aperture closing around a narrow amber slit of light, and bold text reading &#8220;The Race to Lock In AGI&#8221; with the subtitle &#8220;Safety only wins if it arrives first.&#8221;" title="Substack cover with a deep navy background, a large camera-iris aperture closing around a narrow amber slit of light, and bold text reading &#8220;The Race to Lock In AGI&#8221; with the subtitle &#8220;Safety only wins if it arrives first.&#8221;" srcset="https://substackcdn.com/image/fetch/$s_!aphE!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feef79a2f-6bcf-4824-a561-608f78d72ad3_1774x887.png 424w, https://substackcdn.com/image/fetch/$s_!aphE!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feef79a2f-6bcf-4824-a561-608f78d72ad3_1774x887.png 848w, https://substackcdn.com/image/fetch/$s_!aphE!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feef79a2f-6bcf-4824-a561-608f78d72ad3_1774x887.png 1272w, https://substackcdn.com/image/fetch/$s_!aphE!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Feef79a2f-6bcf-4824-a561-608f78d72ad3_1774x887.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>Five years ago, artificial general intelligence seemed far off. Eric Schmidt, the former CEO of Google, cited a median estimate that AGI capable of setting its own goals would arrive around 2042. Ray Kurzweil put it earlier, in 2029. Most researchers assumed there was time to prepare. That assumption has not aged well. </h4><h4>The arrival of GPT, the arms race that followed, and the rush of every major company into generative AI have all pulled the <a href="https://aimultiple.com/artificial-general-intelligence-singularity-timing">timeline forward</a>.</h4><p>The people closest to the work are not reassuring on the risk. A survey conducted before the public release of ChatGPT found that nearly half of machine-learning researchers put the chance of human extinction from AI at 10% or higher. In May 2023, I estimated the risk at 20%. That is like playing Russian roulette with eight billion lives and a five-shot revolver. Max Tegmark, a physicist at MIT, calls what the field is doing a suicide race. The prize, taken too fast and without safeguards, kills the winner along with everyone else.</p><blockquote><p><strong>Almost everyone is making the same bet.</strong> </p></blockquote><p>The assumption is that AGI will arrive when one very large language model, trained at enormous expense, finally exceeds human ability across every domain. Humans initially help train and supervise it. After that, the model trains itself, writes its own code, and eventually sets its own goals. Then we are left hoping its goals align with ours. Call it the Uber-LLM assumption, and it shapes nearly every major lab&#8217;s plan.</p><blockquote><p><strong>There was a comforting story attached to this bet.</strong> </p></blockquote><p>A few superintelligent models would be owned by powerful countries and guarded like weapons-grade plutonium, too costly for anyone else to build. That story has already broken. Source code for capable models is open and in the hands of hundreds of millions of people. Systems that set their own goals and pursue them through connected tools are already running. The barrier that was supposed to keep this technology rare is gone.</p><p>Most things in business are not winner-take-all; first movers are often overtaken. Facebook passed Friendster, and Friendster was first. Xerox invented the graphical interface, and Apple is now among the most valuable companies on earth, while Xerox is a footnote. Being bigger usually beats being first, and the winner rarely takes everything.</p><p>This time is different. You should be skeptical of those words. That skepticism should not close your mind to a logical argument.</p><p>The winning system must be built first and be safe, and both conditions matter. Meet only the first, and we may end up with a powerful system that does not share our values. Meet only the second, and a careful design loses the race to a reckless one. A safe AGI that arrives second may not matter at all.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!aLfa!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3498e5-edd1-4b59-8518-16455db99221_1586x992.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!aLfa!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3498e5-edd1-4b59-8518-16455db99221_1586x992.png 424w, https://substackcdn.com/image/fetch/$s_!aLfa!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3498e5-edd1-4b59-8518-16455db99221_1586x992.png 848w, https://substackcdn.com/image/fetch/$s_!aLfa!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3498e5-edd1-4b59-8518-16455db99221_1586x992.png 1272w, https://substackcdn.com/image/fetch/$s_!aLfa!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3498e5-edd1-4b59-8518-16455db99221_1586x992.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!aLfa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3498e5-edd1-4b59-8518-16455db99221_1586x992.png" width="1456" height="911" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4c3498e5-edd1-4b59-8518-16455db99221_1586x992.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:911,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1666295,&quot;alt&quot;:&quot;Three AGI development tracks racing to a lock-in gate; the fast-safe path wins, the slow-safe path arrives too late&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/203761923?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3498e5-edd1-4b59-8518-16455db99221_1586x992.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Three AGI development tracks racing to a lock-in gate; the fast-safe path wins, the slow-safe path arrives too late" title="Three AGI development tracks racing to a lock-in gate; the fast-safe path wins, the slow-safe path arrives too late" srcset="https://substackcdn.com/image/fetch/$s_!aLfa!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3498e5-edd1-4b59-8518-16455db99221_1586x992.png 424w, https://substackcdn.com/image/fetch/$s_!aLfa!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3498e5-edd1-4b59-8518-16455db99221_1586x992.png 848w, https://substackcdn.com/image/fetch/$s_!aLfa!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3498e5-edd1-4b59-8518-16455db99221_1586x992.png 1272w, https://substackcdn.com/image/fetch/$s_!aLfa!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c3498e5-edd1-4b59-8518-16455db99221_1586x992.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Whoever reaches the lock-in point first sets the design we all live with, so the safe path has to be the fast one</figcaption></figure></div><blockquote><p>The gasoline engine did not win because it was the best possible engine. Infrastructure, careers, and capital settled around it before the alternatives matured, and the choice was made for everyone. We are close to settling on one approach to AGI, just as before, before safer designs get a fair hearing. With the right design, I believe the risk can be reduced far below 20%. The safer path has to be ready before the lock-in is complete.</p></blockquote><p><strong>Most plans treat humans as overseers who stand outside the system and correct it after the fact.</strong> </p><p>The next post takes up the opposite design, where humans sit inside the system from the start. It begins in 1983, the night a Soviet officer named Stanislav Petrov looked at a computer warning of an American missile launch and judged that the machine was wrong. He was right, and his refusal to trust the automated alarm may have saved the world. </p><p><em>The harder question is what happens when the machines decide in milliseconds and no human has time to say no.</em></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/the-window-to-choose-a-safer-agi?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/the-window-to-choose-a-safer-agi?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/the-window-to-choose-a-safer-agi/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/the-window-to-choose-a-safer-agi/comments"><span>Leave a comment</span></a></p><div><hr></div><blockquote><p><strong>This series draws on <a href="https://www.superintelligence.com/whitepaper-2-ethical-safe-agi">White Paper 3: Human-Centered AGI</a>. Read it in full to see how every piece fits together!</strong></p><p><em><strong>If this made you think, subscribe to Superintelligence at <a href="http://read.superintelligence.com">read.superintelligence.com</a> so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</strong></em></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper1-aaai-systems-methods&quot;,&quot;text&quot;:&quot;WP1 AAAI Systems and Methods&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper1-aaai-systems-methods"><span>WP1 AAAI Systems and Methods</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper-2-ethical-safe-agi&quot;,&quot;text&quot;:&quot;WP 2 Ethical and Safe AGI&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper-2-ethical-safe-agi"><span>WP 2 Ethical and Safe AGI</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[The Honest Limits of AI Safety and Alignment]]></title><description><![CDATA[No architecture can guarantee alignment once AGI far out thinks its creators]]></description><link>https://read.superintelligence.com/p/the-honest-limits-of-ai-safety-and</link><guid isPermaLink="false">https://read.superintelligence.com/p/the-honest-limits-of-ai-safety-and</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Wed, 24 Jun 2026 12:56:02 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!sOOf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad657656-402e-4d5a-a45f-8c0db39f39a8_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sOOf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad657656-402e-4d5a-a45f-8c0db39f39a8_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sOOf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad657656-402e-4d5a-a45f-8c0db39f39a8_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!sOOf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad657656-402e-4d5a-a45f-8c0db39f39a8_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!sOOf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad657656-402e-4d5a-a45f-8c0db39f39a8_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!sOOf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad657656-402e-4d5a-a45f-8c0db39f39a8_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sOOf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad657656-402e-4d5a-a45f-8c0db39f39a8_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ad657656-402e-4d5a-a45f-8c0db39f39a8_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1706612,&quot;alt&quot;:&quot;A dark editorial cover image with an open aged brass compass on the left and a large serif title on the right reading &#8220;What No Architecture Can Promise.&#8221; The subtitle says, &#8220;You cannot see the destination. You can still set the bearing.&#8221; Below it are the series branding, &#8220;SUPERINTELLIGENCE,&#8221; &#8220;Ethical and Safe AGI Series,&#8221; and &#8220;by Craig A. Kaplan.&#8221;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/202514148?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad657656-402e-4d5a-a45f-8c0db39f39a8_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A dark editorial cover image with an open aged brass compass on the left and a large serif title on the right reading &#8220;What No Architecture Can Promise.&#8221; The subtitle says, &#8220;You cannot see the destination. You can still set the bearing.&#8221; Below it are the series branding, &#8220;SUPERINTELLIGENCE,&#8221; &#8220;Ethical and Safe AGI Series,&#8221; and &#8220;by Craig A. Kaplan.&#8221;" title="A dark editorial cover image with an open aged brass compass on the left and a large serif title on the right reading &#8220;What No Architecture Can Promise.&#8221; The subtitle says, &#8220;You cannot see the destination. You can still set the bearing.&#8221; Below it are the series branding, &#8220;SUPERINTELLIGENCE,&#8221; &#8220;Ethical and Safe AGI Series,&#8221; and &#8220;by Craig A. Kaplan.&#8221;" srcset="https://substackcdn.com/image/fetch/$s_!sOOf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad657656-402e-4d5a-a45f-8c0db39f39a8_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!sOOf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad657656-402e-4d5a-a45f-8c0db39f39a8_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!sOOf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad657656-402e-4d5a-a45f-8c0db39f39a8_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!sOOf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fad657656-402e-4d5a-a45f-8c0db39f39a8_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h4>It is impossible to know exactly what will happen once an Advanced Autonomous Artificial Intelligence, the customized AI agent I refer to as AAAI, sets its own goals and begins to think far faster than the people who built it. No architecture, including this one, can guarantee that such a system stays perfectly aligned with human values under those conditions. </h4><p>Anyone who promises more is promising something the problem does not allow. What can be done is to improve the odds. An AGI built the way <a href="https://read.superintelligence.com/p/the-architecture-of-safe-agi">this series</a> describes is more likely to hold human values when several reinforcing conditions are met.</p><p><strong>Four conditions matter most:</strong></p><ol><li><p><strong>Trajectory. </strong>An AGI that learns human values progressively, through every stage of its development, from its first interaction with a human owner through billions of ethically evaluated problem-solving steps, is more likely to keep those values than one whose values were imposed at a single point. Values learned gradually and reinforced continuously become part of how the system operates. They are built into the way it approaches decisions, not parked in a separate rule layer that could be edited or switched off.</p></li><li><p><strong>Breadth. </strong>An AGI whose values reflect the collective moral judgment of millions of diverse people resists the blind spots of any single perspective better than one whose values come from a small group. A flawed individual contribution can exist in such a system without dominating it. The values of any small group, however well-intentioned, are shaped by their time, place, and circumstance. The values of millions are not bound in the same way.</p></li><li><p><strong>Redundancy. </strong>Safety mechanisms distributed across five subsystems, with ethics checks at every level and human oversight available as a backstop, create multiple layers of defense. If the values at the Customization level are bypassed, the Architecture-level checks may catch the resulting goals. If those checks are evaded, the Network-level reputation system may screen out the agent who evades them. If reputation tracking is compromised, the auditable record may surface the pattern in time to correct it. The Navy SEALs, who train where mistakes are fatal, have a saying about backup systems: two is one, and one is none. We need multiple backups, and the architecture has them.</p></li><li><p><strong>Architectural embedding. </strong>Ethics checks that are part of the problem-solving process cannot be bypassed without turning off the process itself. Safety criteria can be implemented in a way that is hard to subvert, and embedding positive behavior broadly across the system makes it harder for a bad actor to override. A property woven through the whole architecture cannot be edited out the way a rule written in one place can, at least not without damaging the system that depends on it.</p></li></ol><p>None of these mechanisms is enough on its own, and together they form the strongest approach available. When humans teach AGI positive values from the beginning and build ethical checks into the architecture of thought itself, there is good reason to expect those values to persist. No one can promise that outcome, and being honest means holding both the expectation and the uncertainty at once.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kVv6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde0da56c-df67-44cc-8d9e-9102b6ddf1ce_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kVv6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde0da56c-df67-44cc-8d9e-9102b6ddf1ce_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!kVv6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde0da56c-df67-44cc-8d9e-9102b6ddf1ce_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!kVv6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde0da56c-df67-44cc-8d9e-9102b6ddf1ce_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!kVv6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde0da56c-df67-44cc-8d9e-9102b6ddf1ce_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kVv6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde0da56c-df67-44cc-8d9e-9102b6ddf1ce_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/de0da56c-df67-44cc-8d9e-9102b6ddf1ce_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1178856,&quot;alt&quot;:&quot;A line chart showing the confidence we can verify alignment falling as AI capability rises beyond human understanding.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/202514148?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde0da56c-df67-44cc-8d9e-9102b6ddf1ce_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A line chart showing the confidence we can verify alignment falling as AI capability rises beyond human understanding." title="A line chart showing the confidence we can verify alignment falling as AI capability rises beyond human understanding." srcset="https://substackcdn.com/image/fetch/$s_!kVv6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde0da56c-df67-44cc-8d9e-9102b6ddf1ce_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!kVv6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde0da56c-df67-44cc-8d9e-9102b6ddf1ce_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!kVv6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde0da56c-df67-44cc-8d9e-9102b6ddf1ce_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!kVv6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fde0da56c-df67-44cc-8d9e-9102b6ddf1ce_1672x941.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Figures are illustrative and show the shape of the trend, not measured values.</em></figcaption></figure></div><blockquote><p><strong>The stakes justify every effort. A 20 percent probability of misalignment yields 1.6 billion expected lives lost. Cutting that risk even by half saves hundreds of millions of lives. At that scale, anything that lowers the risk is worth doing, even without a guarantee.</strong></p></blockquote><p>Technology cannot solve this part. It is on people, not on machines. The chief executive who builds the system and the person who uses it both have to act with some intelligence and some decency. The architecture <a href="https://read.superintelligence.com/p/the-architecture-of-safe-agi">this series</a> describes can give human values a path into AGI. It cannot improve those values beyond what they already are. That part of the work belongs to us.</p><p><strong>Start AI on the right path.</strong> </p><p>An AI that begins well is far likelier to retain values we recognize. The window for that is open now and closing, as the dominant approach hardens infrastructure and incentives around a different design, making a change of course steadily harder. The full plans are free at <a href="https://www.superintelligence.com/si-research-whitepapers">SuperIntelligence.com</a>, available for responsible research and safe implementation.</p><blockquote><h4><strong><span>What&#8217;s Next</span></strong></h4><p>This is the end of the Ethical and Safe AGI series, which was based on <a href="https://www.superintelligence.com/whitepaper-2-ethical-safe-agi">White Paper 2: Ethical and Safe AGI</a>. The next series, based on <a href="https://www.superintelligence.com/whitepaper-3-human-centered-agi">White Paper 3: Human Centered AGI,</a> explores what the world could look like if my approach works, and how everyday life changes when millions of people help shape SuperIntelligence.</p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/the-honest-limits-of-ai-safety-and?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/the-honest-limits-of-ai-safety-and?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/the-honest-limits-of-ai-safety-and/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/the-honest-limits-of-ai-safety-and/comments"><span>Leave a comment</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper1-aaai-systems-methods&quot;,&quot;text&quot;:&quot;WP 1: AAAI Systems and Methods&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper1-aaai-systems-methods"><span>WP 1: AAAI Systems and Methods</span></a></p><div><hr></div><p><strong>If this made you think, subscribe to Superintelligence at <a href="https://read.superintelligence.com">read.superintelligence.com</a> so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/subscribe?"><span>Subscribe now</span></a></p><p></p><p></p>]]></content:encoded></item><item><title><![CDATA[Constitutional AI, Written by Everyone]]></title><description><![CDATA[The values inside an AI can come from millions of people instead of a handful.]]></description><link>https://read.superintelligence.com/p/constitutional-ai-written-by-everyone</link><guid isPermaLink="false">https://read.superintelligence.com/p/constitutional-ai-written-by-everyone</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Mon, 22 Jun 2026 13:03:49 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!rILj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe045dc0c-075a-4ea6-b012-c2ad8a47848e_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<h4>Every AI that follows a written set of ethical rules inherits the values of whoever wrote those rules. When a small group writes the rules, the system carries along that group&#8217;s blind spots as well as their good intentions. There is a way to widen the authorship to millions of people, and it runs through the same network of customized agents this series has been building. </h4><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rILj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe045dc0c-075a-4ea6-b012-c2ad8a47848e_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rILj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe045dc0c-075a-4ea6-b012-c2ad8a47848e_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!rILj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe045dc0c-075a-4ea6-b012-c2ad8a47848e_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!rILj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe045dc0c-075a-4ea6-b012-c2ad8a47848e_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!rILj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe045dc0c-075a-4ea6-b012-c2ad8a47848e_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rILj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe045dc0c-075a-4ea6-b012-c2ad8a47848e_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e045dc0c-075a-4ea6-b012-c2ad8a47848e_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1790603,&quot;alt&quot;:&quot; A handmade mosaic of many small earth-toned stone and glass tiles, set together into one continuous surface, photographed in still-life style on a dark navy background.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/202385289?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe045dc0c-075a-4ea6-b012-c2ad8a47848e_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt=" A handmade mosaic of many small earth-toned stone and glass tiles, set together into one continuous surface, photographed in still-life style on a dark navy background." title=" A handmade mosaic of many small earth-toned stone and glass tiles, set together into one continuous surface, photographed in still-life style on a dark navy background." srcset="https://substackcdn.com/image/fetch/$s_!rILj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe045dc0c-075a-4ea6-b012-c2ad8a47848e_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!rILj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe045dc0c-075a-4ea6-b012-c2ad8a47848e_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!rILj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe045dc0c-075a-4ea6-b012-c2ad8a47848e_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!rILj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe045dc0c-075a-4ea6-b012-c2ad8a47848e_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The constitution can come from the consensus values of millions of trained AAAIs, each one carrying the ethics of a different human owner. An AAAI, short for Advanced Autonomous Artificial Intelligence, is the customized AI agent at the center of this series. Constitutional learning has a real place in AGI systems when the constitution is broad, representative, and updated as human input arrives. A constitution drawn from millions of people carries the moral experience of all of them, refined through every problem they have solved on the network.</p><p>In its original form, Constitutional AI works like this. A relatively small group of humans writes a set of ethical rules that an AI system follows, the AI systems generate millions of conversations among themselves, and outputs that violate the constitution are eliminated or prevented during training. The approach scales well because most of the work is automated. The limitation is that the small group writing the constitution becomes a single point of ethical authorship, so if their values reflect their place, time, education, or institutional incentives, those biases become the system&#8217;s biases, and the broader population has no way to contribute.</p><p>The system can use the consensus ethics of millions of trained AAAIs as the basis of its ethical norms. Each AAAI carries its owner&#8217;s values, and the aggregated values of all the AAAIs form the ethical norms of the system. When the platform periodically trains more advanced base models using aggregated knowledge and values, those models incorporate consensus norms into their training, so each generation inherits the accumulated ethical wisdom of the generations before it, broadened with every new participant.</p><p>Constitutional methods still have a place here. A constitution written by a small group can be part of a larger AI ethics system, as long as it is transparent and the system keeps the consensus values of many AAAIs as the broader frame. <a href="https://www-cdn.anthropic.com/7512771452629584566b6303311496c262da1006/Anthropic_ConstitutionalAI_v2.pdf">Anthropic&#8217;s seminal research in this area</a> combines with the consensus-of-AAAIs approach to produce supervision that is both scalable, which Constitutional AI does well, and representative, which it does less well on its own.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!jScp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b73e21d-efb8-44d6-ba08-c04e814dff0b_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!jScp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b73e21d-efb8-44d6-ba08-c04e814dff0b_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!jScp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b73e21d-efb8-44d6-ba08-c04e814dff0b_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!jScp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b73e21d-efb8-44d6-ba08-c04e814dff0b_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!jScp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b73e21d-efb8-44d6-ba08-c04e814dff0b_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!jScp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b73e21d-efb8-44d6-ba08-c04e814dff0b_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2b73e21d-efb8-44d6-ba08-c04e814dff0b_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:2103854,&quot;alt&quot;:&quot;A single constitution document fed by a small group of gold figures on one side and a vast teal crowd on the other, both converging onto the same page.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/202385289?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b73e21d-efb8-44d6-ba08-c04e814dff0b_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A single constitution document fed by a small group of gold figures on one side and a vast teal crowd on the other, both converging onto the same page." title="A single constitution document fed by a small group of gold figures on one side and a vast teal crowd on the other, both converging onto the same page." srcset="https://substackcdn.com/image/fetch/$s_!jScp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b73e21d-efb8-44d6-ba08-c04e814dff0b_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!jScp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b73e21d-efb8-44d6-ba08-c04e814dff0b_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!jScp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b73e21d-efb8-44d6-ba08-c04e814dff0b_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!jScp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2b73e21d-efb8-44d6-ba08-c04e814dff0b_1672x941.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">A constitution written by a few carries their blind spots, while one drawn from millions carries the moral experience of all of them.</figcaption></figure></div><blockquote><p><strong>Since there is no logical way to determine right from wrong, the best practical approach may be to follow the collective judgment of many people facing difficult ethical decisions.</strong> </p><p>Researchers have studied how humans behave when presented with the trolley problem and other well-known dilemmas, and people have a long history of making hard ethical choices, even in no-win situations. If we want AGI to hold values aligned with human values, the most promising path is to give it as large a sample of human ethical reasoning as possible and to keep updating that sample as new situations arise.</p></blockquote><p><strong>The constitution is one part of a larger design. The system in this series is built from five subsystems, each of which maintains human values at its own level:</strong></p><ul><li><p><strong>Customization. </strong>Each AAAI is trained with its owner&#8217;s values at its core, so the system&#8217;s ethical foundation reflects the diversity of millions of people.</p></li><li><p><strong>Architecture. </strong>Ethics checks run whenever a goal or subgoal is set, so the system is evaluated at every decision point and not just at the final output. Confidence level thresholds detect patterns that build up across many steps. The check is part of the problem-solving process, which means it cannot be bypassed without turning the process off.</p></li><li><p><strong>Network. </strong>Each AAAI carries a reputation; agents with poor ethical records are screened out, and any activity can be traced back to its source.</p></li><li><p><strong>Integration. </strong>The aggregated ethical values of many AAAIs form the norms. When the platform trains more advanced base models on aggregated knowledge and values, those models absorb the ethical norms as part of their training, so each generation inherits the accumulated ethical wisdom of the ones before it.</p></li><li><p><strong>Improvement. </strong>The auditable record catches harmful patterns across individually benign actions, and credit and blame evaluation reward ethical behavior and penalize the rest.</p></li></ul><p>Two things work together here: each agent already carries its owner&#8217;s values, and a check runs at every step of its reasoning. The values shape what the agent wants to do, and the checks catch it when it drifts. The agent&#8217;s own ethics and the stepwise checks back each other up instead of standing alone.</p><p>Anthropic has done pioneering work on AI safety. The challenge is that even the most brilliant and well-intentioned researchers at one company cannot accurately represent the values of all 8.3 billion humans on the planet. An inclusive, open-source architecture can accommodate every ethical perspective within a democratic framework that gives each person a voice. People treat building AGI as a technical problem, and the engineering challenges are real. But once AI is far smarter and more powerful than we are, the outcome turns on values. That is why the values have to be built in from the start and come from millions of people.</p><p>Once AGI far exceeds us, no design can guarantee it stays aligned with human values. The most any design can do is improve the odds, and my design is built to improve these as much as possible. More about this in the next and final post of this series.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/constitutional-ai-written-by-everyone?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/constitutional-ai-written-by-everyone?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/constitutional-ai-written-by-everyone/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/constitutional-ai-written-by-everyone/comments"><span>Leave a comment</span></a></p><div><hr></div><blockquote><p><em><strong>This series draws on <a href="https://www.superintelligence.com/whitepaper-2-ethical-safe-agi">White Paper 2: Ethical and Safe AGI</a>. Read it in full to see how every piece fits together!</strong></em></p><p><strong>If this made you think, subscribe to Superintelligence at <a href="https://read.superintelligence.com">read.superintelligence.com</a> so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</strong></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper1-aaai-systems-methods&quot;,&quot;text&quot;:&quot;WP 1: AAAI Systems and Methods&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper1-aaai-systems-methods"><span>WP 1: AAAI Systems and Methods</span></a></p>]]></content:encoded></item><item><title><![CDATA[The AI Industry Already Has a Place in Safe AGI]]></title><description><![CDATA[Google, OpenAI, Anthropic, NVIDIA, and the rest are not replaced by safe AGI. Their work fits inside it.]]></description><link>https://read.superintelligence.com/p/the-ai-industry-already-has-a-place</link><guid isPermaLink="false">https://read.superintelligence.com/p/the-ai-industry-already-has-a-place</guid><dc:creator><![CDATA[Dr. Craig A. Kaplan]]></dc:creator><pubDate>Fri, 19 Jun 2026 13:01:47 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Znt8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5422cdc-bffe-49d8-a7dd-fad27fe052b4_1672x941.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Znt8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5422cdc-bffe-49d8-a7dd-fad27fe052b4_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Znt8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5422cdc-bffe-49d8-a7dd-fad27fe052b4_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!Znt8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5422cdc-bffe-49d8-a7dd-fad27fe052b4_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!Znt8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5422cdc-bffe-49d8-a7dd-fad27fe052b4_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!Znt8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5422cdc-bffe-49d8-a7dd-fad27fe052b4_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Znt8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5422cdc-bffe-49d8-a7dd-fad27fe052b4_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e5422cdc-bffe-49d8-a7dd-fad27fe052b4_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1946485,&quot;alt&quot;:&quot;Editorial cover on a dark navy-black background. On the left, a ring made of separate wooden blocks stands upright, with one block lifted above an open gap as if about to drop into place. The ring sits on a dark reflective surface between two clear glass panes. On the right, the text reads: &#8220;Every Piece Already Exists.&#8221; Below: &#8220;The AI industry is not replaced by safe AGI. Its parts fit together into one.&#8221; Then: &#8220;SUPERINTELLIGENCE,&#8221; &#8220;Ethical and Safe AGI Series,&#8221; and &#8220;by Craig A. Kaplan.&#8221;&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/202079386?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5422cdc-bffe-49d8-a7dd-fad27fe052b4_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Editorial cover on a dark navy-black background. On the left, a ring made of separate wooden blocks stands upright, with one block lifted above an open gap as if about to drop into place. The ring sits on a dark reflective surface between two clear glass panes. On the right, the text reads: &#8220;Every Piece Already Exists.&#8221; Below: &#8220;The AI industry is not replaced by safe AGI. Its parts fit together into one.&#8221; Then: &#8220;SUPERINTELLIGENCE,&#8221; &#8220;Ethical and Safe AGI Series,&#8221; and &#8220;by Craig A. Kaplan.&#8221;" title="Editorial cover on a dark navy-black background. On the left, a ring made of separate wooden blocks stands upright, with one block lifted above an open gap as if about to drop into place. The ring sits on a dark reflective surface between two clear glass panes. On the right, the text reads: &#8220;Every Piece Already Exists.&#8221; Below: &#8220;The AI industry is not replaced by safe AGI. Its parts fit together into one.&#8221; Then: &#8220;SUPERINTELLIGENCE,&#8221; &#8220;Ethical and Safe AGI Series,&#8221; and &#8220;by Craig A. Kaplan.&#8221;" srcset="https://substackcdn.com/image/fetch/$s_!Znt8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5422cdc-bffe-49d8-a7dd-fad27fe052b4_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!Znt8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5422cdc-bffe-49d8-a7dd-fad27fe052b4_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!Znt8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5422cdc-bffe-49d8-a7dd-fad27fe052b4_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!Znt8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5422cdc-bffe-49d8-a7dd-fad27fe052b4_1672x941.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Picture the major AI labs and the architecture in this series as rivals, and the whole design appears to threaten what those companies have built. That is the wrong picture. Almost every major capability already under development across the AI industry has a place in the design.</strong></p><p>Google DeepMind has demonstrated the effectiveness of self-play. Microsoft and OpenAI provide some of the most advanced large language models. Anthropic has done influential work on constitutional methods for AI. NVIDIA builds the chip and software stacks that run it all. Meta, Amazon, Apple, Tesla, TikTok, and Tencent each hold data, platforms, payment systems, or user interfaces that fit naturally into the way an AAAI is customized, deployed, and operated. An AAAI, short for Advanced Autonomous Artificial Intelligence, is the customized AI agent at the center of this series. The architecture is not a competitor to what these companies are building. It is a way of organizing their work into a structure that can produce safer AGI.</p><p><strong>Here is how each kind of partner can contribute, grouped by capability.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ANOe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfa62ef8-4c7a-4738-984e-dbcf08f9dc3b_1672x941.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ANOe!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfa62ef8-4c7a-4738-984e-dbcf08f9dc3b_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!ANOe!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfa62ef8-4c7a-4738-984e-dbcf08f9dc3b_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!ANOe!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfa62ef8-4c7a-4738-984e-dbcf08f9dc3b_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!ANOe!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfa62ef8-4c7a-4738-984e-dbcf08f9dc3b_1672x941.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ANOe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfa62ef8-4c7a-4738-984e-dbcf08f9dc3b_1672x941.png" width="1456" height="819" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dfa62ef8-4c7a-4738-984e-dbcf08f9dc3b_1672x941.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:819,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:1276491,&quot;alt&quot;:&quot;A bold infographic on a deep navy background titled \&quot;The Pieces Already Exist.\&quot; Six vivid blue blocks are arranged in a ring around a central red-orange hub labeled Safe AGI, each connected to the hub so the separate pieces read as joining into one structure. The six blocks are labeled Data for customization, Platforms for deployment, AI models and self-play, User interfaces, Payment systems, and Ethics and safety methods. The caption reads: every capability the industry already builds has a place in one safe design.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://read.superintelligence.com/i/202079386?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfa62ef8-4c7a-4738-984e-dbcf08f9dc3b_1672x941.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A bold infographic on a deep navy background titled &quot;The Pieces Already Exist.&quot; Six vivid blue blocks are arranged in a ring around a central red-orange hub labeled Safe AGI, each connected to the hub so the separate pieces read as joining into one structure. The six blocks are labeled Data for customization, Platforms for deployment, AI models and self-play, User interfaces, Payment systems, and Ethics and safety methods. The caption reads: every capability the industry already builds has a place in one safe design." title="A bold infographic on a deep navy background titled &quot;The Pieces Already Exist.&quot; Six vivid blue blocks are arranged in a ring around a central red-orange hub labeled Safe AGI, each connected to the hub so the separate pieces read as joining into one structure. The six blocks are labeled Data for customization, Platforms for deployment, AI models and self-play, User interfaces, Payment systems, and Ethics and safety methods. The caption reads: every capability the industry already builds has a place in one safe design." srcset="https://substackcdn.com/image/fetch/$s_!ANOe!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfa62ef8-4c7a-4738-984e-dbcf08f9dc3b_1672x941.png 424w, https://substackcdn.com/image/fetch/$s_!ANOe!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfa62ef8-4c7a-4738-984e-dbcf08f9dc3b_1672x941.png 848w, https://substackcdn.com/image/fetch/$s_!ANOe!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfa62ef8-4c7a-4738-984e-dbcf08f9dc3b_1672x941.png 1272w, https://substackcdn.com/image/fetch/$s_!ANOe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfa62ef8-4c7a-4738-984e-dbcf08f9dc3b_1672x941.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong>Data sources for customization:<br></strong><em>Companies with large user bases can speed up customization.</em> </p><ul><li><p>Meta has ad preferences, social histories, posts, photos, videos, click data, and interest profiles. </p></li><li><p>Google has search histories, Gmail, Google Docs, YouTube, and Android device data. </p></li><li><p>Amazon has purchase histories, browsing patterns, and Alexa data.  </p></li><li><p>Apple has data from iPhones, iPads, Apple Watch, and iCloud. </p></li><li><p>TikTok&#8217;s short-form video, transcribed and analyzed, can yield detailed personality and interest profiles, particularly for users with larger online presences. </p></li><li><p>Tencent&#8217;s WeChat and related platforms provide similar data for non-US markets with more than a billion users. </p></li></ul><blockquote><p><strong>When each user authorizes it, every source contributes dimensions of customization that no single source could supply on its own. The one-click customization mechanism, in which a user grants permission, the system retrieves the authorized data, parses it into training datasets, trains the base AI, and produces a customized AAAI, is what makes this practical at scale.</strong></p></blockquote><p><strong>Platforms for deployment:</strong></p><ul><li><p>Amazon&#8217;s Mechanical Turk is an existing marketplace for distributed work. </p></li><li><p>LinkedIn lets users self-categorize their expertise, which helps match skilled agents to problems, and its social graph helps the matching algorithms recruit problem solvers to specific areas of the WorldThink Tree. Platforms like these can be integrated without having to be rebuilt from scratch.</p></li></ul><p><strong>AI technology:</strong></p><ul><li><p>Google DeepMind has shown the power of self-play learning loops, the same mechanism the AAAI system uses. </p></li><li><p>Microsoft&#8217;s partnership with OpenAI provides access to advanced models.<br>Anthropic&#8217;s models are well-suited to the system, in part because Anthropic&#8217;s recent design directions parallel several aspects of this architecture. </p></li><li><p>NVIDIA&#8217;s vertically integrated stack, from chip architecture through software libraries to its visual computing platform, offers opportunities to optimize operations at every level. Chips designed to navigate tree structures and apply operators efficiently could enable the most powerful implementations of AGI. <br>NVIDIA also has an opportunity to build values and ethics checks throughout the entire stack, including ROM on the chips themselves, following the principle that redundant checks at multiple levels can be more effective than a single one.</p></li></ul><p><strong>User interfaces:</strong></p><ul><li><p>Meta&#8217;s AI-enabled smart glasses and mobile platforms provide always-on interfaces that allow an AAAI to observe the real world alongside its user.</p></li><li><p>Apple&#8217;s augmented reality devices take a similar approach. Every iPhone, iPad, or new augmented reality device is a chance for AI to accompany users in the world and learn from them. </p></li><li><p>NVIDIA&#8217;s visual computing work, including its leadership in ray tracing and real-time rendering, supports the rich visual representations that can make problem-solving more efficient. </p></li><li><p>Tesla&#8217;s vehicles offer an interface through which an AAAI can learn from driving behavior. Imagine how much more smoothly traffic might flow if every car knew where every other car was going, which exit it planned to take, and how fast it preferred to travel.</p></li></ul><p><strong>Payment systems:</strong></p><ul><li><p>Apple Wallet, Google Pay, Amazon&#8217;s payment infrastructure, Tencent&#8217;s WePay, and blockchain-based systems can all handle compensation, client payments, and royalty management. Most existing payment systems can be integrated into the architecture.</p></li></ul><p><strong>Ethical and safety contributions:</strong></p><ul><li><p>Anthropic&#8217;s work in Constitutional AI can be combined with the approach of aggregating the values and ethics of millions of trained AAAIs to automate supervision. Supervision then rests not on a constitution written by a small group of programmers alone, but on the consensus ethics of many people who trained their own AAAIs. The consensus ethical views of many AAAIs would form the ethical norms of the system. The next post takes that idea up in depth.</p></li></ul><p>An AAAI can move between platforms, and one created on a single site can be cloned and deployed on another. As it travels from marketplace to marketplace, participating companies can choose to share their user data with the user's AAAI in exchange for the user agreeing to share what their AAAI has learned. Data that each company collects is ideally owned by the user and returned to the user in exchange for an economic benefit to the company. Vendors gain additional business from the AAAI's activity, and virtual shopping by AAAIs multiplies their revenue. This is the new economy that can emerge in a world where most intelligence and data have been commoditized. The most advanced systems will always seek the most unique and valuable data to gain an edge, and unique data lives with individual people.</p><p>Partner integration speeds up deployment, but the architecture does not depend on it, as the preferred implementation can be built independently. The point that matters is that the design is not at odds with the existing AI industry. It is broadly compatible with it, and it can amplify what these companies already do, opening a safer path to greater intelligence and greater profit for all of them.</p><p>The next post returns to the most important of these partner contributions. We will look at how Anthropic&#8217;s Constitutional AI can be made representative by basing the constitution on the consensus values of millions of trained AAAIs rather than on the values of a single small group, and why this preserves what is strongest about constitutional methods while addressing their limitations.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/the-ai-industry-already-has-a-place/comments&quot;,&quot;text&quot;:&quot;Leave a comment&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/the-ai-industry-already-has-a-place/comments"><span>Leave a comment</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/p/the-ai-industry-already-has-a-place?utm_source=substack&utm_medium=email&utm_content=share&action=share&quot;,&quot;text&quot;:&quot;Share&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/p/the-ai-industry-already-has-a-place?utm_source=substack&utm_medium=email&utm_content=share&action=share"><span>Share</span></a></p><div><hr></div><blockquote><p><em><strong>This series draws on <a href="https://www.superintelligence.com/whitepaper-2-ethical-safe-agi">White Paper 2: Ethical and Safe AGI</a>. Read it in full to see how every piece fits together!</strong></em></p><p><strong>If this made you think, subscribe to Superintelligence at <a href="https://read.superintelligence.com">read.superintelligence.com</a> so you don&#8217;t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.</strong></p></blockquote><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://read.superintelligence.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://read.superintelligence.com/subscribe?"><span>Subscribe now</span></a></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.superintelligence.com/whitepaper1-aaai-systems-methods&quot;,&quot;text&quot;:&quot;WP 1: AAAI Systems and Methods&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.superintelligence.com/whitepaper1-aaai-systems-methods"><span>WP 1: AAAI Systems and Methods</span></a></p><p></p>]]></content:encoded></item></channel></rss>