Now that Escalade is over and the results are up, we can analyze how accurate the model was. One thing I can say, it certainly is better than my posting predictions, which were actually worse than blind guessing. I should probably stick to statistical probabilities and let others handle the posting divination. One thing that irked me was that the prediction list didn't end up being entirely accurate, some on it did not compete and some debaters (including Adam Wilkins) were not included. I don't think the difference should be too significant though and it will hopefully balance out in the end. So without further ado, let's look at just how accurate the model was by looking at individual prediction buckets:
Of the 66 predictions given a 0-5% chance of occurring, 0 occurred. (Small note, the average of these predictions was about 1.5% so this result isn't particularly surprising)
Of the 50 predictions between 5-15%, 7 occurred.
Of the 30 predictions between 15-25%, 5 occurred.
Of the 30 predictions between 25-35%, 13 occurred.
Of the 39 predictions between 35-45%, 12 occurred.
Of the 10 predictions between 45-55%, 7 occurred.
Of the 11 predictions between 55-65%, 3 occurred.
Of the 9 predictions between 65-75%, 7 occurred.
Of the 4 predictions between 75-85%, 3 occurred.
Of the 7 predictions between 85-95%, 7 occurred.
Here is a graph of the results:
It's not surprising we see it perform more accurately at Escalade than at Ambassador because there are more debaters at Escalade and thus a larger sample size. All of these fell within the 95% confidence interval except for the 55-65% prediction bucket, which as you can see is a big dip on the graph. What happened is that quite a few debaters who were slightly favored not to check beat the odds and checked anyway. The opposite effect happened with the 45-55% prediction bucket and so I'm not too concerned. After all, we should expect some variance, if we have 10 prediction buckets we shouldn't be surprised one of them falls outside the 95% confidence interval. It was probably just a coincidence. And besides, both sample sizes were pretty small. A statistically significant sample size would probably be around at least 30 and we only have that sample size for the 35-45% prediction bucket and everything below that. That's probably why you see less variance at the bottom of the graph. In short, this is a good sign that the model is indeed accurate. You can be confident that if it gives an event a 40% chance at happening, it will happen 40% of the time.
Now let's combine the Ambassador graph and the Escalade graph and we can keep updating it after every tournament. Over time we should start to see the lines draw closer and closer together. I'll post this graph on a separate page and update it over time.
At both of the last tournaments, the 45-55% bucket performed quite a bit higher than expected, and the 55-65% one a bit lower than expected, which is why you see the bumps. But other than those two neighbouring buckets, which if averaged would be right on, the model is extremely close to what is expected. It is near perfect accuracy so far. We have great evidence the model is performing well.
One more quick note on the model's accuracy. It predicted in LD there would be on average 1 6-0, 5 5-1s and 11 4-2s. This is exactly what happened. In TP it predicted there would be on average 1 6-0, 4 5-1s, and 9 4-2s. In reality there were no 6-0s, 7 5-1s, and 7 4-2s.
Of the 66 predictions given a 0-5% chance of occurring, 0 occurred. (Small note, the average of these predictions was about 1.5% so this result isn't particularly surprising)
Of the 50 predictions between 5-15%, 7 occurred.
Of the 30 predictions between 15-25%, 5 occurred.
Of the 30 predictions between 25-35%, 13 occurred.
Of the 39 predictions between 35-45%, 12 occurred.
Of the 10 predictions between 45-55%, 7 occurred.
Of the 11 predictions between 55-65%, 3 occurred.
Of the 9 predictions between 65-75%, 7 occurred.
Of the 4 predictions between 75-85%, 3 occurred.
Of the 7 predictions between 85-95%, 7 occurred.
Here is a graph of the results:
Now let's combine the Ambassador graph and the Escalade graph and we can keep updating it after every tournament. Over time we should start to see the lines draw closer and closer together. I'll post this graph on a separate page and update it over time.
One more quick note on the model's accuracy. It predicted in LD there would be on average 1 6-0, 5 5-1s and 11 4-2s. This is exactly what happened. In TP it predicted there would be on average 1 6-0, 4 5-1s, and 9 4-2s. In reality there were no 6-0s, 7 5-1s, and 7 4-2s.
Comments
Post a Comment