DRIFT DEEP DIVE · PREDICTION, SURPRISE & INTELLIGENCE
Is compression
intelligence?
Better prediction means fewer bits. That connection is precise. The leap to intelligence is not. Watch the deep dive, then make the idea useful with an experiment you can control.
One question. A much bigger connection.
The original narrated Drift exploration, with illustrations and music. Click the picture to play, or jump to a chapter. The film loads only when you ask for it.
CHOOSE A CHAPTER · TAP TO WATCH
Why does guessing help shrink a file?
Predictable things need less explanation.
Imagine sending AAAAAA. If both sides know the rule, “six A's” can describe the pattern more compactly than spelling it out. The rule and the count still have a cost. For a very short file, overhead can erase the saving.
Lossless means nothing is discarded. Saying “store only the mistakes” is a helpful intuition, not a literal description of every compressor. The decoder needs enough information to recover every symbol.
For probabilities ½, ¼ and ¼, the prefix codes A = 0, B = 10, C = 11 use 1.5 bits per symbol on average, instead of a fixed two bits. Common events get short codes; rare events cost more. This is the core idea behind Shannon's source coding theory (1948).
How can a code cost a fraction of a bit?
Not by cutting one physical bit in half. Encode a long sequence jointly and divide its total length by the number of symbols. Arithmetic coding progressively narrows an interval according to the model's probabilities.
In the film, P(A) = 0.6 and P(B) = 0.4 put the two-symbol message AB in [0.36, 0.60). Its width is 0.24, giving an ideal information cost of −log₂(0.24) ≈ 2.06 bits. The decoder must share the model and know the message length or an end marker. A real finite code also needs termination and precision overhead. The coding construction ↗
Same coin. Different model. Different cost.
Suppose independent tosses land heads 90% of the time. A fair-coin model spends one ideal bit per toss. A model that gets the bias right spends about 0.469 bits per toss on average. The coin has not changed. The description has improved.
How much does a wrong guess cost?
Reality stays fixed at 90% heads. Change only the model's prediction.
A fair-coin model spends 1.000 bit per toss. Matching the bias reduces the average cost to 0.469.
The total is cross-entropy: how costly real data is under your model. The unavoidable part is entropy. The extra cost of the mismatch is KL divergence. Think of it as the price of a mistaken model, not a generic “distance between minds.” The information-theory foundations ↗
The equations, if you want them
Let p = 0.9 be the true heads probability and q be the model's guess. All logarithms here use base two.
- Entropy, the best possible average
- H(p) = −p log₂ p − (1−p) log₂(1−p)
- Cross-entropy, your model's average
- H(p,q) = −p log₂ q − (1−p) log₂(1−q)
- KL divergence, the extra cost
- D(p∥q) = H(p,q) − H(p)
A probability of 1/1,000 has surprise log₂(1,000) ≈ 9.97 bits. Assigning zero probability to an event that happens gives infinite ideal log loss. Real compressors use smoothing or fallback codes rather than trying to write an infinitely long file.
Why this matters for language models.
Their prediction loss can become a code length.
Next-token cross-entropy rewards assigning high probability to the token that actually appears. Feed those probabilities into a suitable coder, and lower loss means fewer ideal bits on that evaluated data. Training losses often use natural logarithms; divide by ln(2) to express them in bits.
The foundational 2023 paper, revised in 2024, Language Modeling Is Compression demonstrated lossless image and audio compression with a primarily text-trained model. But its data-only results exclude the model's weights. Those weights, the decoder and computation are not free. Their cost may be shared across many files, but it must be accounted for in a practical comparison.
A recent test: compression and benchmark performance.
Lower technical-text compression rates were strongly associated with higher zero-shot MMLU accuracy, a multiple-choice knowledge benchmark. That is a relationship across models, not proof that shrinking files causes intelligence.
*Spearman ρ; reported 95% interval [−0.942, −0.783]. Read the result ↗
The authors use newly published text to reduce contamination risk. They explicitly say it does not eliminate that risk and does not capture post-training instruction-following abilities. A compression score is informative, not a complete report card for an assistant. Read the limitations ↗
A good map is not the whole journey.
Compression rewards finding regularities. To act intelligently, a system must also use what it knows: choose among actions, respond to feedback and handle a changed situation. For an engineering decision, test those abilities directly instead of replacing them with one file-size number.
The Hutter Prize's official rationale connects compressing human knowledge with prediction and intelligence. That is a serious research motivation. It is not evidence that a winning archive has demonstrated every capacity of a general-purpose agent.
Better compression can reveal better prediction. Whether that becomes useful intelligence depends on what the system can do with the prediction, and how you test it.
Four questions that cut through an AI claim.
- What was preserved?Lossless encoding recovers the exact data. A summary drops detail. Do not confuse a smaller prompt with a lossless archive.
- What did the score leave out?Ask about model size, compute, latency, context length and whether the test data was genuinely unseen.
- What happens when the pattern changes?A predictor tuned to yesterday's distribution may become expensive or unreliable tomorrow.
- Does it help with the actual goal?For an agent, test task completion, errors and recovery. For your notes, test whether you can reconstruct the argument and its caveats.
This matters now: an August 6, 2026 preliminary study found that repeated agent-context compression could weaken recent interactions and increase blocked actions or repeated exploration. That is lossy summarization, a different problem from encoding a file exactly. Smaller context is not automatically better performance.
Choosing the right metric is part of a larger choice: what are you actually trying to improve? Private consulting helps connect the AI landscape to your priorities and direction.
ONE QUESTION CAN OPEN A WHOLE MAP
Try the next branch.
- Go Deeper.“Why does cross-entropy measure prediction quality?”
- Wander.“How is a compressed file different from a scientific theory?”
- Test the edge.“When can a better predictor still make a worse decision?”
The recorded exploration reached 28 ideas and seven levels. Those are map counts, not 28 verified facts. Follow the branches that improve your understanding, then check the claims that matter.
Explore your own question, free ↗No account or API key needed. AI can be wrong; leave out personal or confidential information.
The evidence, in one place.
Checked September 30, 2026. The original film is preserved. This companion adds qualifications: ideal bit costs are not actual file sizes, model weights count, and compression is not a complete intelligence test.
Sources and what they support
- Shannon, 1948: information entropy and source coding. A dated mathematical foundation.
- MacKay, Information Theory, Inference, and Learning Algorithms, author draft (2002): surprise, coding, cross-entropy and relative entropy. University-hosted mathematical foundations. The coin example is calculated here, not an empirical finding.
- Delétang et al., 2023, revised 2024: predictive models as lossless compressors, cross-modal experiments, and model-size accounting. Not a latest-model benchmark.
- Tan et al., September 23, 2026: 80 base models, 14 text categories and the technical-text/MMLU association. A preprint with explicit contamination and instruction-following limitations.
- Hutter Prize official FAQ: the contest's compression/intelligence rationale, not proof that an archive is an autonomous agent.
- Min et al., August 6, 2026: preliminary evidence about lossy agent-context compaction and execution instability.
General education, not professional advice. The coin graphic is a mathematical illustration. Research counts and correlation are labeled separately.
FROM UNDERSTANDING TO PRACTICE
Keep the curiosity.
Give it a direction.
Ready to make this a practice? Compare the Get Amplified and 1000x Lab options on Patreon, where the current tier details are explained.
Explore membership on Patreon ↗


