One bit at a time
Claude Shannon defined information, gave it a unit, and proved its limits. Modern AI still runs on what he found.
You have certainly used the word “information” today. You sent some, received some, maybe complained about having too much of it. And if someone asked you to define it, to say precisely what information is in a way that would let you measure it the way you measure weight or temperature, you would probably pause.
Everyone pauses. For most of human history, information was treated like love or beauty: obviously real, obviously important, impossible to quantify.
Then, in 1948, a 32-year-old Bell Labs researcher named Claude Shannon published a paper that ended the pause. He defined information, gave it a unit, proved theorems about its limits, and in doing so founded the field on which telephones, the internet, data compression, and modern artificial intelligence all rest.
This issue is about that idea, and about the unusual mind that produced it. Because the way Shannon worked is at least as instructive as what he built.
A serious mind that refused to be serious
Claude Elwood Shannon was born in 1916 in a small Michigan town. As a boy he built model planes, a telegraph line to a friend’s house, and gadgets from whatever was around.
He never really stopped. At MIT and later Bell Labs, colleagues knew him as the man who juggled while riding a unicycle down the hallway, who built a mechanical mouse named Theseus that could learn its way through a maze, and who constructed a box whose only function was to switch itself off: flip the switch on, a mechanical hand emerges, flips the switch off, and retreats.
It would be easy to write these off as amusing eccentricities and nothing more. That would miss the point. Shannon treated serious problems the way other people treat puzzles, with curiosity instead of reverence, and with a willingness to look foolish.
In the early 1960s he and mathematician Edward Thorp built what is now considered the first wearable computer and took it to Las Vegas to predict roulette outcomes, with their wives along as lookouts. Not for the money. Because the problem was interesting.
His biographers Jimmy Soni and Rob Goodman argue that this playfulness was inseparable from his productivity: a mind unafraid of “trivial” problems is a mind free to notice connections that status-conscious minds walk past.
One of those connections changed everything.
The master’s thesis that built the digital world
In 1937, Shannon was a 21-year-old graduate student at MIT, working part-time on an early analog computer. His job involved the machine’s tangle of electromechanical relays, switches that are either open or closed.
Shannon had studied symbolic logic as an undergraduate, including the algebra George Boole had published in 1854: a system for calculating with just two values, true and false, using the operations AND, OR, and NOT. For 80 years, Boole’s algebra had interested logicians and philosophers but had found no practical application.
Shannon saw what nobody had seen: a relay circuit and a Boolean expression are the same object. A switch is a truth value. Switches in series compute AND. Switches in parallel compute OR. Which means circuit design, until then a craftsman’s art of intuition and tinkering, could be done with algebra. You could prove a circuit correct. You could simplify it on paper before building it.
He wrote it up as his master’s thesis, “A Symbolic Analysis of Relay and Switching Circuits”. It has been called the most important master’s thesis of the twentieth century, and the description is hard to argue with: it is the founding document of digital circuit design. Every processor since is, at bottom, Boolean algebra made of matter, today executing trillions of Boole’s operations per second to run, among other things, the neural networks reshaping your industry.
Notice the shape of the achievement. Shannon did not invent new field in mathematics in 1937. He noticed that mathematics which already existed, and had been dismissed as useless for 83 years, described a technology that its inventor could never have imagined. The insight was an act of translation, not creation. Hold onto that; it is the recurring pattern of this entire series.
Information is surprise
Eleven years later, at Bell Labs, Shannon published “A Mathematical Theory of Communication.” The telephone company had a practical problem: engineers were building ever-larger networks to transmit “information” without any rigorous definition of what they were transmitting or any theory of how much a channel could carry.
Photograph of two men working on lines and equipment and three telephone operators at work.
Shannon’s move was radical: he threw away meaning. Whether a message is profound or trivial, true or false, is irrelevant to the engineering problem. What matters is uncertainty. A message carries information precisely to the extent that it tells you something you did not already know.
From this one decision, everything follows. If a message resolves a choice between two equally likely alternatives (heads or tails, yes or no, one or zero), it carries one unit of information. Shannon called that unit the bit, short for binary digit (he credited the coinage to his colleague John Tukey). A message that tells you something you already knew, such as “the sun rose this morning”, carries almost no information at all, however many words it uses.
He then defined the entropy of a source: the average surprise of the messages it produces. A fair coin has 1 bit of entropy per flip, the maximum unpredictability for two outcomes. A two-headed coin has 0: the outcome is certain, so learning it teaches you nothing. English text sits in between. Because letters and words are partly predictable from context, Shannon estimated its entropy at roughly one bit per letter, far below the capacity of the alphabet, which is why text compresses so well.
And then came the theorems. Shannon proved there is a hard mathematical floor beneath compression: no scheme, however clever, can compress a source below its entropy without losing information. And he proved something almost paradoxical about noise. Over any noisy channel, error-free communication is possible up to a computable rate, the channel capacity, if you encode cleverly enough. He proved such codes must exist without showing how to build them, and that challenge would ignite decades of work, including Richard Hamming’s.
Information is the resolution of uncertainty.
The core idea of Shannon’s 1948 framework, in one sentence
Why your loss function is a Shannon measurement
If you work anywhere near machine learning, you use Shannon’s mathematics daily, possibly without knowing its name.
Cross-entropy loss, the quantity most classifiers and language models minimize during training, is Shannon’s framework applied to prediction. It measures, in bits, how surprised your model is by reality. A model that assigns high probability to what actually happens has low cross-entropy. A confidently wrong model has enormous cross-entropy. Training a classifier is, quite literally, minimizing surprise.
Perplexity, the standard evaluation metric for language models, is entropy in disguise. It is two raised to the entropy, to be precise. When a lab reports that a new model has lower perplexity, it is reporting progress toward Shannon’s limit for predicting human text. He sharpened that limit in 1951 with a parlor game, asking people to guess the next letter of a hidden sentence. Every large language model is, in a precise sense, a machine built to win Shannon’s guessing game.
Mutual information, how much knowing one variable reduces uncertainty about another, drives feature selection, representation learning, and much of modern interpretability research. And every ZIP file on your devices lives in the space between raw data and Shannon’s entropy floor, the hard limit on lossless compression. JPEG and MP3 go further still. They throw away the detail you were never going to miss.
It is also widely reported that Anthropic named its AI assistant Claude in his honor, a fitting tribute from a field that runs on his mathematics.
The cruelest irony, and the durable lesson
Shannon spent his later years in the grip of Alzheimer’s disease. The man who gave the information age its foundations gradually lost access to his own. He died in February 2001, without fully grasping the scale of what he had set in motion. The internet then blooming around the world was built, at every layer, on his two papers.
There is grief in that. There is also clarity. Shannon’s ideas did not need their author’s awareness, his promotion, or his fame to keep working. They spread because the foundation was valuable. That is the property that separates foundational work from fashionable work: it compounds without your involvement.
And his method is available to anyone. Shannon was not brilliant because he knew more mathematics than everyone else. Many contemporaries knew more. He was brilliant because he kept carrying the mathematics he knew into rooms where it had never been, and asking one modest question: does this apply here? In 1937 the answer connected a Victorian logician to the digital circuit. In 1948 it connected probability theory to the telephone wire. Each connection looked small. Each one compounded for the better part of a century.
Further references can be found in:
A Mind at Play: How Claude Shannon Invented the Information Age
Jimmy Soni & Rob Goodman · Simon & Schuster, 2017
The definitive Shannon biography and the primary source for the personal details in this issue: the juggling, Theseus, the Las Vegas expedition with Thorp, and his final years.
A Mathematical Theory of Communication
Claude E. Shannon · Bell System Technical Journal, Vol. 27, 1948
The founding paper of information theory. Sections 1 and 2 are readable with modest mathematical background and reward the effort; the definition of entropy and the bit appear in the opening pages.
Symbolic Analysis of Relay and Switching Circuits
Claude E. Shannon · MIT master’s thesis, 1937 (published 1938, Transactions of the AIEE)
The thesis connecting Boolean algebra to circuit design, the founding document of digital logic and the subject of this issue’s first half.
Prediction and Entropy of Printed English
Claude E. Shannon · Bell System Technical Journal, Vol. 30, 1951
The letter-guessing experiments estimating the entropy of English, the direct intellectual ancestor of language-model perplexity evaluation.
The Information: A History, a Theory, a Flood
James Gleick · Pantheon Books, 2011
The popular history book that places Shannon in the longer arc from talking drums to quantum information, the best single companion volume for this entire series.
P.S. Thank you for reading this far.
That alone puts you in rare company: most people scroll past anything that asks them to think slowly.
The fact that you’re here suggests you already value the kind of thinking this newsletter is about. Share it with someone who values it too. Those people are worth finding.
Comment below and share the newsletter/ issue if you think it is relevant!
Feel free to also follow me on LinkedIn (very active), Instagram(not so active) or X (not so active).
Until next week’s issue, keep learning, keep building, and keep thinking like a mathematician.
-Terezija









Hi Terezija,
I absolutely loved the way you think and the way you put your ideas into words. Reading your essays is one of those rare experiences where you keep nodding along, thinking, "Exactly!"
At one point, I even thought, "Maybe I married the wrong woman..." 😄
Of course, I mean that in the most innocent and complimentary sense possible. What I really mean is that your intellectual curiosity, clarity, and way of connecting ideas are genuinely captivating. It's refreshing to find writing that makes you stop, think, and smile at the same time.
Thank you for sharing your thoughts. Looking forward to reading the next one.
Very informative! Thank you for the interesting and information-filled essay. Keep it up!