Usefulness is paramount. In 1948, the mathematician Claude Shannon showed that the less probable an event is, the more information its occurrence carries, thereby making rarity the standard measure of information. That body of work is known as Classical Information Theory. But a dataset can be full of rare and surprising material and still be useless, just as an idea can be new without being any good. Useful information that is already known has little value either, and that is where novelty and rarity come back in. Estimating what a piece of information is worth means weighing both at once.
The approach I have called Kaplan Information Theory, or KIT, starts with the observation that there would be no information without differences. In KIT, goal-relatedness is the dimension that does most of that work. The more closely a piece of information relates to a particular goal, the more valuable it is to the entity pursuing that goal. If it contains the exact solution, its goal-related value is as high as possible.
Where Classical Information Theory treats information as an absolute quantity measured against a probability distribution, goal-related value is always relative to an agent and to what that agent is trying to do. The part that matters most is that information can have quite low Shannon entropy and still be intensely goal-related.
For example, if the goal is to build a fire in the woods without matches or another fire source, information on fire-making using only materials found in the woods would have high goal-relatedness. Information about art history would have low goal-relatedness. The problem solver would rather have common knowledge about fire-starting than scarce knowledge about art history. Here, and generally, goal-relatedness trumps Shannon-sense information value or absolute rarity.
In all conceptions of problem solving, the problem solver has goals. One of the most basic heuristics for achieving goals is Means-Ends Analysis. In Means-Ends Analysis, the problem solver examines the gap between the current problem state and the goal state and tries to apply an operator to reduce or bridge the gap. To apply the heuristic, the problem solver must have a way to determine which operator to use. Just as every intelligent entity has evaluation functions for choosing what brings it closer to its goals, it can have evaluation functions for judging how goal-related a particular piece of information is.
One way to think of this is to imagine an AGI or SuperIntelligence with a single goal, let us say, to extract maximum profits from the financial markets. For such an entity, facing potential datasets to pursue and limited resources, it must choose the datasets that will best help it achieve its goal. It may already have learned so much about financial markets that a new financial dataset contains relatively little information in the Shannon sense, since most of it is predictable from what it already knows. That dataset can still carry far more goal-related information than a dataset on art history, even if the art history dataset would score much higher on surprise.
Shannon entropy measures, although widely used and treated as the main way of thinking about information, are a crude approach, used only when goal information is not present. Without any information about an entity’s goals, pursuing datasets with high Shannon entropy values makes sense. But if the goal is known, it immediately becomes more essential to find goal-related information rather than just unusual or unexpected information.
Goal-relatedness does not operate alone. Suppose an intelligent entity already knows a hundred ways to start a fire in the woods. The value of learning one more is far less than it would be for an entity with the same goal that knew nothing about the subject. Once a goal is specified, the value of a piece of information depends on its goal-relatedness and on what the entity already knows that is also goal-related.
Which is why the art history case can eventually reverse. At some point, everything that can be discovered about machine learning will have been found. If there are huge diminishing returns to finding even a very slightly unusual new piece of information about machine learning, and if the AI had a goal of learning everything, it would eventually focus on art history. If the AI knows nothing about machine learning, the time when it focuses on art history may be far away. If the AI knows almost everything about machine learning and nothing about art history, it will look at art history sooner.
Now that we have developed some intuitions and provided examples showing how KIT differs from Classical Information Theory, it is worth listing other dimensions of difference with practical implications.
At the highest level, difference is the key concept in KIT.
Cost. As an AI learns more about a subject, new information becomes rarer and harder to find, making it more expensive to acquire. Practical intelligence needs a cost model to weigh one rare and expensive piece of information against two less rare and cheaper ones.
Rate of change. One dataset holds historical weather patterns. Another is updated daily. A third updates every hundred milliseconds. Their current contents might be identical yet worth very different amounts because one goes stale far faster than the other.
Perceivability. Events too fast, too slow, too small, or too large to be detected through an entity’s senses or instruments carry no usable information for that entity, whatever they contain in principle. They may carry a great deal for an entity built to perceive them.
Representation. The form in which information is represented changes what an intelligence can perceive, infer, or do with it, which is the recognition buried in the saying that a picture is worth a thousand words.
Context refers to differences not only in the culture, technology, knowledge, goals, representations, and perceptual abilities of a specific intelligent entity, but also in those of other intelligent entities that form its context. Details on making fire, shared with someone who does not know how to make fire, have different value depending on whether that individual alone lacks fire-making knowledge or the entire culture in which the individual lives lacks it. The information is identical. Its value is not.
So the question is not only how new information compares to what a single entity already knows. The entire context and the knowledge of every intelligence the information reaches must be taken into account. Just as the difference between two knowledge bases can be measured, the same can be done across any number of them. Every dimension can be evaluated differently depending on how much context is considered.
Note that this principle applies even to Classical Information Theory. For example, a specific string of characters might appear unusual and contain a large amount of information if compared to just one paragraph of text with no such characters. But if a larger sample is used, one in which the same characters appear frequently and surprise nobody, the assessment changes drastically.
Generally, Information Value can be seen as a function of the dimensions listed above, with different constants weighting the importance of each dimension. Other dimensions of difference may exist, or be discovered, so different functions can be written and optimized to maximize an entity’s intelligence.
Of those dimensions, representation deserves more than a line in a list. Give one AI millions of chess games stored as pictures of the board, and give another the same games written in standard chess notation. Same hardware, same games, and one of them will play far better than the other.
The next post explains why the form of knowledge that takes can raise intelligence sharply without adding a single unit of computing power, and why the players who talk about the Ruy Lopez are thinking about more of the board than the players who talk about moving the bishop.
This series draws on White Paper 6: Catalysts for Growth of SuperIntelligence. Read it in full to see how every piece fits together!
If this made you think, subscribe to Superintelligence at read.superintelligence.com so you don’t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.




