There would be no information without differences. Humans, or any intelligent entity, could not perceive a world unless we could perceive and draw distinctions between this and that. The notion of difference, and the quantification of differences, is the essence of meaningful information. The approach I have called Kaplan Information Theory, or KIT, starts there.
That matters for AGI immediately. If an AI knows nothing you do not already know, and can do nothing you could not already do, it adds no information for you. Its value to you begins with the ways it differs from you.
Observing that differences matter is not by itself new. What KIT adds is the choice of unit. Traditional approaches to Information Theory take a purely mathematical view, estimating the probability of events that cannot be well predicted from known information. KIT deals with goals, states, operators, and functions as the primary relevant information units rather than bits. Where classical approaches discard a vast amount of information, KIT considers higher-level representations that group bits into chunks, chunks into concepts, and concepts into solutions that achieve goals.
The oldest formal account of information fits within KIT as one dimension of difference.
Claude Shannon showed in his 1948 paper, A Mathematical Theory of Communication, that the less probable an event is, the more information its occurrence carries, so an unexpected observation departs further from what a probability model led us to expect.
That body of work has come to be called Classical Information Theory, and the formulation made perfect sense for the problem at hand. Shannon was working out how to send information efficiently and reliably over copper wires from a sender to a receiver, which was his problem at Bell Labs. In that context, it was a rigorous definition of information, with practical implications for a channel's capacity to carry it.
Shannon also said plainly that the meaning of a message was irrelevant to the engineering problem he was solving.
However, differences in expectation are only one type of difference that can be measured.
We commonly say an event carries information if it is news, meaning it was previously unknown to a particular recipient, even if it is not generally surprising. Something is new to me and carries information. The same thing is well known to you and carries no weight. There is a relative aspect of information here that is not explicitly part of the classical theory. Assuming a different probability distribution of expected events for each entity could solve this problem, but that seems cumbersome.
Length raises a related difficulty. The amount of information is sometimes proportional to the number of words in a message, but there are situations in which fewer words convey more. Blaise Pascal wrote in 1657 that he had made a letter longer only because he had not had the time to make it shorter, implying that fewer words would have conveyed more useful information. Shannon’s theory has a great deal to say about redundancy and compression, which is what the longer letter contains. It was never designed to determine whether the shorter letter is more useful to the reader.
KIT considers any difference between two events, datasets, categories, or informational units to be a valid measure of information content. Distinguishable events, objects, or categories of information only exist to the degree that differences exist. An infinite string of 1s contains no information. An infinite string of 0s contains no information. Zero has meaning only if 1 is possible and 1 sometimes exists. Similarly, 1 has meaning only if 0 is a possibility and 0 sometimes exists.
Seeing a 1 after an incredibly long sequence of 0s carries much information, not just because it is unexpected, but because it is finally a difference!
Thinking of information as a measure of difference is more general than considering information as a measure of surprise. Surprise is one type of difference. Any difference, even a non-surprising one, contains information.
Consider two datasets. Where the two sets intersect, there is no new information. The sum of the non-overlapping areas, known in set theory as the Symmetric Difference, represents the new information the datasets contain relative to each other. Picture the two datasets, A and B, as a pair of overlapping circles. The region where the circles overlap is what both datasets already have in common. Everything outside that overlap, in either circle, is the Symmetric Difference.
An example makes it concrete. AI #1 knows everything in the Encyclopedia Britannica. AI #2 knows everything in Wikipedia. AI #3 is the combined knowledge of AI #1 and AI #2. The intersection between Wikipedia and Britannica represents what both AIs know and contains no new information for either. The Symmetric Difference, namely the knowledge in Britannica and not in Wikipedia, plus the knowledge in Wikipedia and not in Britannica, represents the new knowledge of AI #3. Time is not relevant in this example. The new information can be calculated by comparing the static information across the two datasets.
Now consider the same two AIs, except that this time each continues to add to its knowledge. The calculation must account for the static encyclopedias and for any new information added over time, so the intersection and the Symmetric Difference are constantly changing.
The information in both examples is still a matter of difference. In the second case, the difference lies between two sets of information that continue to change, rather than between what was expected and what was observed. For information that can reasonably be represented as sets of comparable units, the Symmetric Difference gives KIT one workable measure of how much two entities differ. For any two intelligent entities, the relevant measure of information is the practical difference between what one entity knows and what the other knows. This can sometimes be operationalized as differences in how the two entities behave, which matters when direct insight into their respective knowledge bases is impossible, or when behavior is more relevant than static knowledge.
That has a consequence worth stating plainly. To the degree that an AI produces nothing a human could not already produce, it offers that human no new information. It may still offer information relative to other humans or other AIs. If there were no differences between two intelligent entities, human or AI, it would be impossible to distinguish one from another.
In the intermediate term, humans care about what AI knows that they do not and how AI behaves differently from how they would. A human teacher has information that a human student lacks, and this information differential lies at the heart of the teaching process. Any intelligent entity can teach another entity something new, human or AI, only if it holds useful information the learner does not already possess.
Some of this may seem obvious, but it has profound implications for measuring information. As AGI becomes more intelligent and superintelligent, the chief concern will be finding new sources of information. That information could be measured in surprisingness, as Shannon suggested. It could also be measured by differences in knowledge bases, in behavior, or in the construction of two entities.
The difference is where the answer starts rather than where it ends. Information consists of differences and, more importantly, useful differences. A dataset can have Shannon information content, differ from everything an AI already knows, and still be useless to it.
The next post sets out the dimensions along which differences can be measured. It explains why the same instructions for making fire carry vastly different amounts of information depending on who else already knows how.
This series draws on White Paper 6: Catalysts for Growth of SuperIntelligence. Read it in full to see how every piece fits together!
If this made you think, subscribe to Superintelligence at read.superintelligence.com so you don’t miss what comes next. And if someone in your life needs to understand where AI is heading, send this to them.




