How AI Can Catch Danger in Real Time
When AI is unsure, it stops and asks.
AI does not need to anticipate every danger in advance to stay safe.
Human and AI agents can dynamically flag potential ethical issues in real time as they are encountered, then present them to other groups of agents for resolution. Rather than relying on experts or crowdsourcing to determine the full range of ethical scenarios in advance, the real-time flagging approach allows AI, AGI, and SuperIntelligent systems to detect potential issues and pause work until additional human input helps the system determine the ethical approach.
Of course, in time-critical situations, pausing or delaying might not always be possible, but the approach can be used for many issues that do not demand an immediate response. Including this dynamic approach of delaying responses until ethical input is received can reduce an otherwise exponential space of possibilities to a manageable size. One implication is that critical, high-stakes issues requiring an immediate response, such as whether to launch a counterattack to a perceived missile launch or other military applications, will require proportionally more path coverage and training in advance than situations where a delay in response is acceptable.
Constitutional AI approaches are generally suboptimal for establishing ethical knowledge bases, partly because they rely on rules developed by an elite group.
However, such approaches might be acceptable as a means of temporarily flagging potentially unethical situations until a representative sample of human ethical judgments can be obtained.
For example, a rule that said an AI can never provide information that might be used to harm other humans might flag potentially dangerous scenarios.
Where possible, responses to such situations could be delayed until they were reviewed by humans or otherwise subjected to deeper review.
Some false positives will occur. Someone might ask about using arsenic to poison rats and have to wait for a response while the AI flags the question and gets other human agents to weigh in on whether answering it, given the context of the conversation, poses a risk to humans. As long as the delay is not too long, it might be acceptable if the delay prevents serious safety issues. Established mathematical methods for determining when a test is doing more harm than good, for example, in the medical profession, can be employed to help quantify these decisions. If we can use AI to determine in real time whether an applicant is a good credit risk, there is no reason that similar algorithms cannot be employed to delay or avoid answering certain potentially dangerous questions.
Ideally, there would be a method for rapid review and appeal of the potentially dangerous cases.

One approach, which minimizes delay to a fraction of a second while still providing some margin of safety, is to have questionable cases reviewed by multiple AI agents to see if there is consensus among them on the request’s safety. Human agents, working more slowly, could override the AI agents (teaching the AIs in the process) upon appeal or when they can get to the prioritized list of issues. Automated means for tracking the frequency and potential impact of unanticipated safety issues could help optimize the use of human decision-making for the most common and most important issues.
These same techniques can improve accuracy even when no safety risk is involved.
One preferred method for reducing hallucinations from LLMs is to have multiple AI agents process the same question and then take the consensus or majority answer as the most correct. This approach might employ versions of the same LLM with different parameter settings to generate multiple responses, or completely different LLM models. Users can set the degree of reliability they desire and are willing to pay for. That choice determines how much redundant processing is performed in the final answer.
The difference between employing imperfect real-time detection of safety issues using existing well-known approaches and doing nothing is huge. The nuances of balancing the opportunity costs of not responding to perfectly harmless questions with the costs of preventing disasters can be refined over time, ideally using a data-driven approach. However, real-time detection and prevention of issues before they occur is almost certainly a net positive, even at some threshold of false positives.
All of this depends on AI acquiring human knowledge and values in the first place.
From a user interface perspective, one of the simplest methods is for humans to have conversations with the AI they are customizing and then provide instructions to that AI on how it should behave when training other AIs. Such conversations can be initiated by either the AI being customized, the human doing the customization, or both. Humans with strong beliefs or knowledge about certain issues may want to focus conversations and subsequent AI customization in these areas.
In addition to conversing with humans, AI can conduct surveys to elicit their opinions and knowledge on a wide variety of subjects, including ethical views. Survey approaches have the advantage of being well-suited to gathering random, representative samples of human knowledge, using a variety of established online and offline survey methodologies. Like intelligent conversational approaches, where AI can guide the direction and content of the conversation to fill in knowledge gaps, survey methods can also target specific knowledge gaps, including gaps in coverage of certain ethical situations.
Both conversational and survey methods require humans to engage with AI to teach it actively.
However, humans have limited time to engage in such activities, and AI has an almost insatiable appetite for new knowledge. Therefore, AI will have to rely extensively on passive methods of knowledge acquisition, such as are currently employed in the creation of today’s LLMs. Any method that uses the passive digital footprints left by humans as they perform tasks, including online navigation, selecting products and websites, solving problems, and communicating with other humans, can be used to train AI and acquire knowledge.
Knowledge does not stand still. Some of what AI learns changes by the day, while the most fundamental human values change slowly, if at all.
The next post in this series examines the rate of change of knowledge, pictured as a spinning wheel with passing fashions out at the rim and the deepest human values at the motionless center.
This series draws on White Paper 4: Safe, Scalable Artificial General Intelligence. Read it in full to see how every piece fits together!
If this made you think, subscribe to Superintelligence at read.superintelligence.com so you don’t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.



