transcribe

Measuring the entropy of English

3Blue1Brown · 2m · transcribed Jul 2026
More from 3Blue1Brown Business
𝕏 Share ▶ YouTube 📥 PDF 🤖 .md

Section Insights

# 0:00

Claude Shannon's Early Experiments on Text Predictability

How did Claude Shannon begin to understand the compressibility of English text?

Shannon started by analyzing short sequences of characters in English text to track the probabilities of what characters typically followed one another. He realized that while this method worked for short sequences, it became ineffective for longer sequences.

  • Shannon focused on the predictability of English text to understand its compressibility.
  • He used statistical analysis of character sequences to build predictive models.
  • Short sequences are less effective for predicting longer sequences.
# 0:28

The Role of Context in Text Predictability

Why did Shannon need a different approach to understand language compressibility?

Shannon recognized that longer sequences of text provide more context, making predictions more accurate and text more compressible. To explore this, he turned to his wife for a practical experiment.

  • Longer sequences of text offer better context for predictions.
  • Shannon's wife served as a model for human prediction in his experiments.
  • Understanding compressibility required innovative approaches beyond simple statistics.
# 0:57

Shannon's Experiment with Human Prediction

What method did Shannon use to gather data on letter prediction?

Shannon conducted an experiment where he asked his wife to predict letters from a passage. He recorded her correct and incorrect guesses to analyze the information conveyed by the text.

  • Shannon's experiment involved tracking human guesses to understand information transmission.
  • He aimed to create a string of text that maintained the same information with fewer letters.
  • This approach laid the groundwork for his formal definition of information.
# 1:26

Refining the Experiment with Multiple Participants

How did Shannon improve his understanding of letter prediction probabilities?

In his 1950 paper, Shannon expanded his experiment to include multiple participants, recording not just correct or incorrect guesses but also the number of attempts needed to guess the correct letter. This allowed him to estimate implicit probabilities of letter predictions.

  • Shannon's refined experiment involved multiple guessers to enhance data accuracy.
  • He combined statistics with guessing attempts to estimate letter probabilities.
  • This method provided deeper insights into the predictability of English text.
# 1:55

Estimation of English Text Compressibility

What was Shannon's final estimate for the compressibility of English text?

Shannon estimated that English text could be compressed to around 1 bit per character, given sufficient context. This required moving beyond simple data analysis to a more intelligent approach.

  • Shannon's estimate highlighted the potential for significant text compression.
  • Achieving this level of compression requires intelligent modeling rather than just data probing.
  • The concept of compression as a form of intelligence is a key theme in Shannon's work.

Transcript

0:00 Claude Shannon figured out how the compressibility of English text depends on how predictable it is, so he made multiple attempts at understanding the probability of each new character in English. His earliest experiments involved looking at specific short sequences of characters and tracking the statistics of what tended to follow. For example, you could scan through a couple books, look at every instance where you see the letters TH and build up a table of what tends to follow and how often.

0:25 The problem here is that this completely breaks down for longer sequences of characters, most notably those that never show up in the text you're looking at. But longer sequences give more context to guide a prediction, which means that's when text is at its most predictable and hence most compressible. So to really understand the compressibility of language, he needed some other way to ask about these probabilities. So he turned to one of the most readily available models of language available to him in the 1940s: His wife, Betty.

0:53 As the story goes, he pulled out a book and asked her to predict each new letter from a given passage. Every time she guessed incorrectly, he wrote down the correct letter, and every time she guessed correctly, he would replace the letter with a dash. His idea was that this new string of text had fewer actual letters, but it carried the same information, in the sense that it should give just the right prompting for a duplicate of his wife to fill in the entire text.

1:16 This was a start, but his formal definition of information requires knowing the actual probabilities assigned to new letters, so he needed something better. In his 1950 paper, Prediction and Entropy of Printed English, Shannon outlined an experiment interviewing more people where instead of just logging whether their guess was correct or wrong, he would record how many guesses were necessary for his human guessers to come up with the correct next letter. Separately, he combined the idea of statistics with short sequences with number of guesses to make an estimate at the implicit probabilities that people were assigning to the true next letter, based on their number of guesses.

1:50 His final estimate was that given at least 100 characters of context, English should, at least in principle, be compressible down to around 1 bit per character. And I want you to notice, in order to make this estimate, he was forced to go beyond pure data analysis and probe at something intelligent. Today, more than 75 years later, the way we actually achieve compression close to this limit is not through merely probing at intelligence, but by making our best attempts to engineer it.

2:17 This all comes from the first video in a short series I'm making about the idea that compression is intelligence. If you want more, take a look at the channel.

Summary

Claude Shannon explored the predictability of English text to understand its compressibility, initially using character sequences and later involving human predictions to refine his estimates. His work led to the conclusion that with sufficient context, English text could be compressed to about 1 bit per character, highlighting the interplay between data analysis and human intelligence in understanding language.

- Shannon's early experiments focused on character sequences to track predictability.
- He faced limitations with longer sequences that weren't present in the text.
- To gain better insights, he used his wife as a model for predicting letters in a passage.
- His method involved tracking correct and incorrect guesses to estimate probabilities.
- In his 1950 paper, he recorded the number of guesses needed to predict letters.
- Shannon estimated that with 100 characters of context, English could be compressed to approximately 1 bit per character.
- His work illustrates the need to combine data analysis with human intuition for understanding language compressibility.
- The series suggests that modern compression techniques aim to engineer intelligence rather than just analyze it.

Questions Answered

How did Claude Shannon begin to understand the compressibility of English text?

Shannon started by analyzing short sequences of characters in English text to track the probabilities of what characters typically followed one another. He realized that while this method worked for short sequences, it became ineffective for longer sequences.

Why did Shannon need a different approach to understand language compressibility?

Shannon recognized that longer sequences of text provide more context, making predictions more accurate and text more compressible. To explore this, he turned to his wife for a practical experiment.

What method did Shannon use to gather data on letter prediction?

Shannon conducted an experiment where he asked his wife to predict letters from a passage. He recorded her correct and incorrect guesses to analyze the information conveyed by the text.

How did Shannon improve his understanding of letter prediction probabilities?

In his 1950 paper, Shannon expanded his experiment to include multiple participants, recording not just correct or incorrect guesses but also the number of attempts needed to guess the correct letter. This allowed him to estimate implicit probabilities of letter predictions.

What was Shannon's final estimate for the compressibility of English text?

Shannon estimated that English text could be compressed to around 1 bit per character, given sufficient context. This required moving beyond simple data analysis to a more intelligent approach.

© transcribe · For agents Built with care and craft by Gokul Rajaram