Hacker News new | ask | show | jobs
by mcphage 13 days ago
How does this paper not even mention the word "overfitting"?
1 comments

The abstract does say

> limitations, principally the small number of tournaments available for validation and the risk of in-sample weight selection

But I agree this model is no more valuable than Paul the Octopus.

That almost makes it worse—like they're vaguely aware that training too heavily on too small a data set makes badly trained models, but are unaware that it has a name and is an actual identified problem.
They've even described how they overfitted it! For five world cups, a simple model based on ranking and goal difference in the group stages[1] predicted it four times, so they invented a somewhat subjective variable on defensive strength to cover the teams that didn't score much...

[1]yes, both of those are endogenous variables...

I appreciate everyone's feedback. This is a preprint, and I was looking exactly for this kind of constructive scrutiny to help improve the work. The data is limited, and I'm well aware of that. I just thought this would be a good first step, with encouraging results (and, after today's match, accurate ones too).