Skip to content
echohive
← All field notes

DRIFT DEEP DIVE · LEARNING THE PATTERN, NOT JUST THE EXAMPLES

Why doesn’t giant AI
always overfit?

Big models can memorize. They can also learn patterns that carry beyond their training data. The difference depends on the data, the learning process, and how we test the result.

· Film 5:58 · About an 8 minute read

Why Giant AI Doesn’t Overfit

The original Drift film, with narration, illustrations and music. Tap the picture to play, or choose a chapter. The reading companion adds qualifications: giant models sometimes do overfit, and the examples below are not AI benchmark results.

Prefer to read? ↓

CHOOSE A CHAPTER · TAP TO WATCH

01 / THE APPARENT CONTRADICTION▶ Watch 0:00

Enough room to memorize everything.

Imagine teaching a model to recognize cats. With real labels, different pictures share useful structure. Replace those labels with random answers and the same powerful model can still fit the training set. But an arbitrary label on a new picture is not predictable from the picture.

The random-label experiments published at ICLR 2017 made this puzzle concrete: convolutional networks could fit nonsense labels, not just meaningful ones. Capacity alone therefore cannot explain why useful training generalizes.

MEMORIZATION“I have seen this.”

Reproduce a training example or its particular answer.

GENERALIZATION“The pattern carries over.”

Predict well on genuinely unseen examples from the relevant distribution.

These are not opposites. One model can do both. Overfitting is the harmful part: fitting training-specific detail or noise in a way that worsens performance beyond that training data.

Data matters before any clever theory does. Varied examples can reveal a relationship worth reusing. Arbitrary labels cannot. A pile of near-duplicates is not the same as broad coverage of the situations where the model will be used.

02 / DOUBLE DESCENT▶ Watch 0:43

Bigger can hurt, then help.

The familiar story is a U-shaped curve: too little flexibility misses the pattern; too much starts fitting noise. Some modern models show a second descent. Around the point where a model can first fit all the training data, test error can spike. Beyond that point, error can fall again.

A schematic double-descent curveTest error falls, rises around the interpolation threshold, then falls again as model capacity increases. This is a conceptual diagram, not measured data. Test errorModel capacity → Too little flexibilityError can fall againCan first fit all training data
Schematic, not benchmark data. The height, location and even presence of the peak depend on the setting.

Nakkiran and colleagues’ 2019 study found double descent across several deep-learning setups, including changes in model size and training duration. It does not establish that every larger model performs better, or that today’s language models all lie on this curve.

Useful distinction

Double descent describes a pattern in test error. It is not, by itself, the mechanism that explains it.

03 / WHEN AN EXACT FIT CAN BE BENIGN▶ Watch 1:16

Fitting noise need not ruin every prediction.

In the film’s house-price example, floor area carries substantial signal while many other measurements carry little. Under suitable conditions, an exact fit can distribute training noise across many weak directions rather than distort the important predictive directions.

Bartlett and colleagues’ linear-regression theory identifies conditions under which a minimum-norm exact fit predicts well despite fitting noisy observations. The data’s covariance structure matters, including having sufficiently many prediction-unimportant directions relative to the sample size.

The conditions do the work. Extra features are not automatically harmless. If important signal lives in weak directions, or the test distribution shifts, the reassuring story may fail.

04 / WHICH PERFECT FIT?▶ Watch 1:43

A training algorithm has preferences.

Suppose one training example only demands w₁ + w₂ = 1. Many weight pairs fit it perfectly. Their training error cannot tell them apart.

Two exact fits, different weight sizes
WeightsTraining fitSquared norm
(0.5, 0.5)Exact0.5
(1, 0)Exact1.0

For this symmetric least-squares problem, gradient descent starting at zero keeps the two weights equal and approaches (0.5, 0.5). More generally, zero-start gradient descent in underdetermined linear least squares selects the minimum-Euclidean-norm solution when it converges to an exact fit.

This is implicit regularization: the algorithm favors a particular solution even without an explicit penalty in the objective. It does not mean every neural network finds a minimum-norm solution.

The catch: “small” depends on the coordinates

Rescale the first feature by 100. The fitting equation becomes 100w₁ + w₂ = 1. Its minimum-norm weights are (100/10001, 1/10001), not (0.5, 0.5). The same Euclidean rule now favors a different prediction function. “Small weights” is meaningful only relative to the representation and units, and does not guarantee the right answer.

05 / TRAINING TIME IS A FILTER▶ Watch 2:50

Not every direction learns at the same speed.

In linear least squares, some directions of the data are learned quickly and others slowly. Stopping early can leave the slow ones mostly untouched. If they contain noise, that helps. If they contain important signal, it hurts.

STRONG DIRECTION99.9%fitted after 10 steps
10× WEAKER SINGULAR VALUE4.9%fitted after the same 10 steps

A controlled linear example from the film, not measured language-model performance. Step size 0.5; squared singular values 1 and 0.01.

Show the tiny calculation

For a least-squares gradient-descent direction with squared singular value λ, a stable step size η and a zero start, the fitted fraction after t steps is f(t) = 1 − (1 − ηλ)ᵗ. The example gives 1 − 0.5¹⁰ = 99.9023% and 1 − 0.995¹⁰ = 4.8890%. A 10× difference in singular value becomes a 100× difference in squared singular value. The film’s earlier 100× singular-value example becomes 10,000× in this learning-rate scale, leaving 4.88% fitted after 1,000 steps.

That is why early stopping is more than “use a smaller model.” It changes how much of each direction is learned. There are rigorous results for particular regression settings, including Raskutti, Wainwright and Yu’s 2014 analysis. They are not universal guarantees for deep networks.

TRY IT / A ONE-DIRECTION EXPERIMENT▶ Watch 3:47

When does learning more make things worse?

Drag the slider to learn more of a noisy estimate. First try a weak signal. Then switch to a useful, clearer signal. The right amount of fitting changes completely.

Signal strength squared: 0.04. Noise variance: 0.25.

Signal left unlearned0.0144
Noise let into the fit0.0400
Total excess error0.0544
How bias and variance change as more of a direction is fittedThe interactive plot shows unlearned signal, admitted noise and their total. Numeric values are displayed above and explained below.Share fitted →0%100%Excess error
Unlearned signalAdmitted noiseTotalLowest error

At 40%, total excess error is 0.0544. For this noisy direction, the theoretical minimum is at 13.8% fitted.

A toy shrinkage calculation, not a training simulator or an AI benchmark. “40% fitted” means the share learned in one direction, not 40% of training time. Error is in illustrative squared-prediction units; unavoidable new-observation noise is not included.

What is the experiment calculating?

Suppose the estimated coefficient equals its true value β plus zero-mean noise of variance τ². Retain a fraction f of that estimate. Expected squared coefficient error is R(f) = (1 − f)²β² + f²τ². The first term is squared bias; the second is variance. For a unit-variance feature direction, this is also excess prediction error. The optimum is f = β² / (β² + τ²). The slider does not derive f from epochs. Real training must infer a stopping point without knowing the true signal.

The film’s rounded numbers, about 0.05 versus 0.20, become 0.0544 versus 0.2029 using its exact parameters. That is about 3.73× more error. In the strong-signal preset, fitting most of the direction is better. Early stopping is a trade-off, not a virtue by itself.

06 / WHAT RECENT RESEARCH ADDSRead the sources ↓

The theory is useful. It is not the whole story.

Linear examples isolate mechanisms beautifully. Modern language and image models involve changing representations, nonlinear optimization, different objectives and many kinds of training. Three recent papers help keep the explanation honest:

REVISED SEPTEMBER 27, 2026 · PREPRINT

Remembering is not the same as using.

A fine-tuning study by Dai and colleagues found a gap between recalling newly injected facts and using them in multi-hop reasoning. Their internal-representation interventions support a possible routing explanation in the tested models. This is not a complete explanation of LLM generalization, but it is a sharp warning against treating recall scores as understanding.

POSTED JULY 6, 2026 · WORKSHOP PAPER

Repeating data can still help.

Zuo and colleagues’ data-reuse study reports preliminary settings where repeating high-quality data keeps improving downstream performance beyond a commonly cited four-epoch limit. A full adaptive scheduler remains future work. There is no universal rule that another pass through the data must be harmful.

JULY 2, 2026 · PREPRINT

Some model families need a different explanation.

Farghly and colleagues’ diffusion-model analysis finds that benign overfitting and double descent do not carry over in the settings they study. Their theory and image-generation experiments instead emphasize mechanisms that prevent harmful fitting. Do not apply a linear-regression story to every AI system.

The calibrated conclusion

We have explanations and proofs for important settings, not one settled theory that predicts every frontier model’s behavior. Data coverage, architecture, explicit regularization and optimization bias all matter.

07 / THE PRACTICAL TEST▶ Watch 4:07

Better training numbers are not enough.

The useful question is: does the improvement survive outside the examples that produced it? A lower training loss, a memorized fact, or a perfect demonstration cannot answer that alone.

  1. Hold out what really needs to be new.Deduplicate near-identical examples. When appropriate, split by person, organization or time, not just by row. A new row from the same source may be an easy leak.
  2. Tune on validation, not on your final test.Use validation to choose checkpoints and settings. Repeatedly reacting to the same test score makes that test part of your decision process. Keep a final evaluation untouched.
  3. Test use, not just recall.Ask for a new combination, an unfamiliar case, or an explanation that requires the learned relationship. Compare with a simple baseline.
  4. Look at both curves and the mistakes.If training keeps improving while validation gets worse, investigate. Repeat with different seeds or splits; one lucky score is weak evidence.
  5. Check the deployment distribution.Good held-out performance can still fail when the users, language, inputs or environment change. Generalization within one distribution is not universal robustness.

This checklist is practical guidance, not a proof that a system is safe. It is a way to make an improvement claim harder to fool yourself with.

Sources, dates and boundaries.

This is a researched reading companion to the supplied Drift film, not a verbatim transcript. The narration and content are unchanged; the video is optimized for web playback. The recent research below was posted or revised within the last three months; older papers are explicitly retained as mathematical foundations, not as current model benchmarks.

Open the seven primary sources
  1. Zhang et al., Understanding deep learning requires rethinking generalization. ICLR 2017, revised February 26, 2017. Foundational random-label experiments.
  2. Nakkiran et al., Deep Double Descent. December 4, 2019. Foundational experiments, not a universal scaling law.
  3. Bartlett et al., Benign Overfitting in Linear Regression. Revised January 29, 2020; PNAS 2020. Foundational theory under specified covariance conditions.
  4. Raskutti, Wainwright and Yu, Early Stopping and Non-parametric Regression. JMLR 2014. Mathematical foundation in kernel regression.
  5. Dai et al., Why Memorized Knowledge Fails to Generalize in LLM Finetuning. Revised September 27, 2026. Preprint; scoped fine-tuning experiments.
  6. Zuo et al., Train Smarter, Not Longer. Posted July 6, 2026; DATA-FM workshop paper. Preliminary data-reuse findings, not a completed universal scheduler.
  7. Farghly et al., Benign Overfitting Does Not Occur in Diffusion Models. July 2, 2026. Preprint; the result has model and distribution assumptions.

The double-descent diagram is schematic. The interactive lab and percentages are deterministic toy calculations. Neither measures a production model’s accuracy. The cover is a conceptual illustration, not a plot of research data. Research checked October 1, 2026.

FOLLOW YOUR OWN QUESTION

A question becomes
a map you can explore.

This film follows a seven-level Drift exploration through 28 ideas. That is a learning path, not 28 verified facts. Try the free browser app, branch into related ideas, and use primary sources to check the claims that matter.

Explore your own question in Drift ↗

Free to try. No account or API key needed. AI explanations can be wrong.

KEEP THE CURIOSITY. STRENGTHEN THE PRACTICE.

Understand more.
Choose what matters.

LEARN AT YOUR OWN PACE

Get Amplified

Move from interesting ideas to a repeatable practice: AI workflows, research methods, markets and attention.

Explore the field guide ↗
THINK IT THROUGH TOGETHER

1000x Lab

Bring your questions to the Sunday conversation. Explore changing ideas and their implications through live discussion and replays.

See how the Lab works ↗
APPLY IT TO YOUR OWN DIRECTION

Private consulting

Step back from the tools. Clarify the larger system, what deserves your attention, and where your next move has leverage.

Explore private consulting ↗

Ready to make this a practice? Compare the Get Amplified and 1000x Lab options on Patreon, where the current tier details are explained.

Explore membership on Patreon ↗