Research Hub
Topic Hub · Memorising isn't learning.

Machine Learning.

Generalisation, leakage and the difference between a model and a description.

Personalized feedUpdated · 2026-07-08

Articles

0

Quick Takes

0

Videos

0Soon

Podcasts

0Soon

Clarity Index

Soon

Followers

Soon

The Standard on Machine Learning

A model that fits the past perfectly has usually memorised it. Machine learning coverage lives or dies on one distinction: performance on data the model has seen, versus data it has not.

What this field actually measures

Train/test split

Holding out data the model never trains on.

ReferenceThe minimum standard; cross-validation is stronger.

Overfitting gap

Difference between training and validation performance.

ReferenceA large gap means the model learned the sample, not the phenomenon.

Feature leakage

Information in the inputs that would not exist at prediction time.

ReferenceThe most common cause of spectacular results that fail in production.

Evidence & metrics

The published work behind the numbers above, and the public datasets you can open to check us. If a claim can’t be traced here, we don’t print it.

Citations

  1. Random forests

    Breiman — Machine Learning · 2001

    Established ensemble averaging as a variance-reduction method and formalised out-of-bag error estimation.

  2. Hidden technical debt in machine learning systems

    Sculley et al. — NeurIPS · 2015

    Model code is a small fraction of a deployed ML system; most failure comes from data and pipeline coupling.

Datasets

  • UCI Machine Learning Repository

    University of California, Irvine · free

    Hundreds of benchmark datasets with documented provenance.

  • OpenML

    OpenML community · free

    Datasets plus reproducible run results for thousands of tasks.

Still open

  • How much sports-modelling accuracy is leakage rather than insight?
  • Which model classes are worth their interpretability cost in decision settings?
  • What should a public accuracy claim be required to disclose?

What we won’t say

We will not quote an accuracy figure without knowing the base rate it beat.

  • VO-1Evidence before wit
  • VO-2Punch at claims, never at people
  • VO-3Say the uncertainty out loud
Read the Brand Voice Standard
Latest Articles

Everything published on this topic.

No articles have been published for Machine Learning yet. New long-form pieces will land here.

The Archive

Everything published on Machine Learning.

The permanent record of every piece we publish under this topic — grouped by year, oldest kept as history.

Quick Takes

Short reactions, polls, and open questions.

Faster than the long form. Same standard.

No quick takes match this filter yet.

2CS Analysis Framework

Evidence, counterpoints, patterns, and the standard.

Every entry links back to its source article. Nothing duplicated.

Framework entries appear here as articles publish.

Future Watch

Open questions and measurable predictions.

We log the question, the confidence, and the review date. When the evidence moves, so does our position.

No active predictions are currently being tracked for this topic.

Video

Coverage on screen.

Long-form, shorts, and playlists connected back to the source articles.

Video coverage is coming soon.

Coming Soon

We haven't published video on Machine Learning yet. Playlists, chapters, and shorts will surface here as they roll out.

Podcast

Long-form conversations.

No podcast episodes yet.

Coming Soon

When we publish podcast episodes on Machine Learning, each will link back to its source article.

Related Topics

Connected hubs.

Where this topic overlaps with the rest of the 2CS knowledge graph.

Topic Discussion

The conversation, on the record.

Community discussion is coming soon.

Coming Soon

A dedicated space for readers to discuss this topic will live here — moderated to the same standard as the editorial.

The 2ND Take

Follow Machine Learning.

One weekly dispatch. Signal over noise. No spam.

No spam. Unsubscribe anytime.