Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Correction--it worked. The authors chose a two-year old sample during which the Dow Jones fell 30.7%. I seriously doubt this will have any predictive power outside of that sample.


Fitting a model to "predict" events that have already happened, isn't anywhere as hard as actually predicting events that have yet to happen.


I'm not so sure about that. Its pretty hard to predict the past too.

Anyone doing serious research into this, will first partition the past data into training and test sets. (And sometimes other validation sets).

So the idea would be to fit a model on one set of past data (the 'training' set), check it works, and then, in the final evaluation, run it on the never seen before, never used, never thought about, 'test' data.

If you have a model trained on 2009, and it also does a great job the first time you run it on the Q1 2010 data that you've never looked at before, I'm now interested, even though every data point is in the past.

I imagine they had to do something like this to pass review.


If you have a model trained on 2009, and it also does a great job the first time you run it on the Q1 2010 data that you've never looked at before, I'm now interested, even though every data point is in the past.

So, you have a model trained on 2009. You try it on the Q1 2010 data and...it doesn't work. Damn. So you throw it out, go back to the drawing board, and try again. And again. And ag...hey, this one works! Trained on 2009 data, and it predicts Q1 2010 perfectly!

Do you trust this model to predict Q2 2010?


Obviously if there has been a 'meta' process of refinement, such that the test set has been used in model development, then its not a clean test set any more, and shouldn't be regarded as such. That's something for any researchers to watch out for, and be sure they don't do. And good researchers are well aware of these pitfalls.

That's why I mentioned the validation sets, and that the test set must never have been looked at, or used before.

But the point stands - if the method works on a clean test set, even if the test set is in the past, then it should be taken seriously.

Would I trust such a model to predict the stock market in Q2 2010? No, because my prior belief is that the stock market is very hard to predict, so I would need very strong evidence to the contrary. But that has nothing to do with having confidence in models that have been tested on historical data.


Right, and they do this for a period from February 2008 to December 2008. We do the same here at the Federal Reserve when developing models.

Sure, the data is historical, but your model doesn't distinguish between "old" and "new" data. If your model predicts test data (in-sample forecasting, right?) well then you have something interesting.


Right, that was my point. It worked for a specially-selected sample during which a bunch of correlated macro factors existed. I commented a bit lower with more details.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: