Results are out from the last tournament, meaning that it's time to analyze whether or not our model performed accurately. It should be noted that this model is not in the business of predicting "winners" but rather assigning probabilities. The model gave Benjamin Lightfoot had the highest probability to win LD - 17% - but it also gave him an 83% chance to lose. So it can't be said that the model was wrong because it didn't predict the winners. The question remains then, how do we analyze how accurate it was? The method I've devised is to make individual prediction buckets (i.e. 0-5%, 5-15% etc). And count what proportion of the time the predictions in those buckets were realized.
One quick disclaimer before moving on, I received incorrect information (one might even call it fake news) regarding which TPers were partnering with whom. This may have affected the overall TP predictions but my guess is that it wouldn't be too significant. I did, however, have to throw out the individual predictions made for the teams I got wrong.
So how did it do? Overall the model made predictions regarding 35 debaters, 20 LD debaters and 15 TP teams. With four predictions made per debater/team (their chance of going 6-0, 5-1, 4-2, and 3-3 below) a total of 140 predictions were made. Here's what happened:
Of the 35 events the model gave a 0-5% chance of occurring, 1 occurred.
Of the 28 events between 5-15%, 2 occurred.
Of the 17 events between 15-25%, 7 occurred.
Of the 17 events between 25-35%, 2 occurred.
Of the 19 events between 35-45%, 8 occurred.
Of the 5 events between 45-55%, 4 occurred.
Of the 7 events between 55-65%, 4 occurred.
Of the 7 events between 65-75%, 3 occurred.
Of the 4 events between 75-85%, 3 occurred.
Of the 1 event between 85-95%, 1 occurred.
Here's a graph displaying the actual results versus the predicted results:
While the model appears to deviate from reality at some points significantly, a general upward trend can be seen, which is a good sign. The reason for the deviation is likely a small sample size and doesn't necessarily imply the model was wrong. For example, the model gave five events approximately a 50% chance of occurring and four of the five ended up occurring. Hence, the large spike in the center of the graph. Does this mean it was wrong? Again, not necessarily. If you flipped a coin five times and it came up heads four of those times, you wouldn't immediately assume the coin was rigged. It's also worth noting that in the probability "buckets" that had a more significant number of entries (i.e. 0-5% which had 35 entries) the model was extremely close to reality. Whether or not that holds true across the board remains to be seen; we'll continue to analyze results from other tournaments when they're complete. Overall, the model performed at the very least decently.
Full results and how they compared to the model's predictions can be found here:
https://docs.google.com/spreadsheets/d/1sJ1F3bQeVAsrgRmGRyItBLffDxUugWGNeS_D8LCal7c/edit?usp=sharing
One quick disclaimer before moving on, I received incorrect information (one might even call it fake news) regarding which TPers were partnering with whom. This may have affected the overall TP predictions but my guess is that it wouldn't be too significant. I did, however, have to throw out the individual predictions made for the teams I got wrong.
So how did it do? Overall the model made predictions regarding 35 debaters, 20 LD debaters and 15 TP teams. With four predictions made per debater/team (their chance of going 6-0, 5-1, 4-2, and 3-3 below) a total of 140 predictions were made. Here's what happened:
Of the 35 events the model gave a 0-5% chance of occurring, 1 occurred.
Of the 28 events between 5-15%, 2 occurred.
Of the 17 events between 15-25%, 7 occurred.
Of the 17 events between 25-35%, 2 occurred.
Of the 19 events between 35-45%, 8 occurred.
Of the 5 events between 45-55%, 4 occurred.
Of the 7 events between 55-65%, 4 occurred.
Of the 7 events between 65-75%, 3 occurred.
Of the 4 events between 75-85%, 3 occurred.
Of the 1 event between 85-95%, 1 occurred.
Here's a graph displaying the actual results versus the predicted results:
While the model appears to deviate from reality at some points significantly, a general upward trend can be seen, which is a good sign. The reason for the deviation is likely a small sample size and doesn't necessarily imply the model was wrong. For example, the model gave five events approximately a 50% chance of occurring and four of the five ended up occurring. Hence, the large spike in the center of the graph. Does this mean it was wrong? Again, not necessarily. If you flipped a coin five times and it came up heads four of those times, you wouldn't immediately assume the coin was rigged. It's also worth noting that in the probability "buckets" that had a more significant number of entries (i.e. 0-5% which had 35 entries) the model was extremely close to reality. Whether or not that holds true across the board remains to be seen; we'll continue to analyze results from other tournaments when they're complete. Overall, the model performed at the very least decently.
Full results and how they compared to the model's predictions can be found here:
https://docs.google.com/spreadsheets/d/1sJ1F3bQeVAsrgRmGRyItBLffDxUugWGNeS_D8LCal7c/edit?usp=sharing
Comments
Post a Comment