Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Wednesday, February 24, 2021

Chance of Precipitation

I heard the term POP on the radio. It stands for "Probability of Precipitation". The explanation of it was pretty simple, but there is a little more to it than I thought. What it measures, of course, is the chance you are going to get rain, snow or whatever, falling on you that day. A fairly simple formula: 

POP = (probability any precipitation falls in the area) x (predicted area of coverage).

Examples: 

The meteorologist thinks about half the region will get wet. There is a 20% chance it rains somewhere in that area. So, 0.50 x 0.20 = 10% chance of rain.

There is a 70% chance of rain falling somewhere in the region. The coverage is almost total, say 90%. So, 0.70 x 0.90 = 63% chance of rain. That is, a 63% chance that it will rain on you.

By the way, that number doesn't tell you anything about how much it might rain. Although by "precipitation" they do seem to hold to there being at least a hundredth of an inch falling.  

I missed out. I think I would have liked to have been a meteorologist.

Monday, January 18, 2021

Wilt's Free Throws

 Wilt Chamberlain is the second leading scorer of all time. Right behind Michael Jordan. But, he couldn't shoot free throws. He once missed 22 straight free throws. That seems worth investigating. First what are the chances? He made 51.1% of his free-throws during his career. In other words he has a 48.9% chance of missing. So to miss two straight would be .489 x .489 = .239121 or a 23.9% chance. Twenty-two in a row? That would be .489 to the 22nd power. That comes out to 0.000000146. 

Just for fun, I took the reciprocal. That is almost seven million. He shot over eleven thousand free throws in his career. I thought missing 22 in a row might be somewhat likely. But no, it isn't. There is a lot of math here a person could play with. Wilt wasn't the worst ever. Andre Drummond makes 38.6% of his free throws. You could figure the probability for him missing 22 in a row.

Could a typical player do this? Lets suppose we want to find out, what percent would you normally have to shoot to have a 50% probability of of missing 22 in a row.  The equation would be x^22 = .50. Using the log and inverse log on your calculator, you come up with 96.9%. That is the chance you miss a shot. So, that means your free throw success rate would be 3.1%. Not good.



Tuesday, January 5, 2021

Election Results

 There was an important Senate election in Georgia (on the website I was looking at it was "Georiga"). I saw this :

Candidate A   51.5%  -  1,957,641

Candidate B   48.5%  -  1,844,815

That was with 87% of the vote counted.

If the vote is split evenly the rest of the way, Candidate A wins, obviously. What percent would be needed? I figured it was my duty as a citizen to find out.

There are a total of 3,802,456 votes so far with 87 % in. That means there must be a total of:

0.87x = 3,802,456

x = 4,370,639

Candidate A needs at least half of that, which is 2,185,319.

With a little subtraction there are 586,183 votes left and he need 227,678 of them. Candidate A needs 40.1% of the remaining votes. That sounds about right. Certain sections get counted at different times. So its not over. 


Update - Its not decided yet, but Candidate A now has 50.02% of the vote. About as close as it can get.

Tuesday, October 6, 2020

Going for Two

Seems like I wrote something like this before, but things have changed now, so it is time to revisit. In the NFL, after you score a touchdown, you can go for a one-point conversion, by kicking, or go for a two-point conversion, by running or passing. The two pointers are harder to make, though. So which should you do? The kicks pretty much always went through. Because of that, the NFL decided they needed to make things a little more interesting and moved the kick back a ways. It still almost always goes through, but the percent is down a little bit - to a success rate of 93.8% of the time. The pass or throw option is roughly half - a success rate of 50.1%.

We should find the expected value for each. 

For a kick, you can get one point, 93.8% of the time:

                    1 x .938 = .938 points per try

For running or passing, you get two points, 47.9% of the time:

                    2 x .501 = 1.02 points per try

Is that even worth messing with? It depends on how good your team is, but typically a team scores something around 60 touchdowns in a season. That would mean roughly (.938 x 60 =) 56.28 points if you kick all the time. And (1.02 x 60 =) 61.2 when running or passing. So about five point over the course of a season. So not a lot, but then again, it only takes one point to lose a game.

And of course your decision could depend on the situation in any particular game. If you score a touchdown near the end of the game and you are now behind by two, you definitely would go for the two-point attempt. 

By the way, I don't know if this was in the book, but Scorecasting is a great little book about sports and how coaches don't always do what makes the most sense. 



Saturday, March 25, 2017

Predicting Win Percentages

Continuing on from last weeks post regarding the website fivethirtyeight.com and how they come up with their information. Last week was about how they look at various political polls and how they rank them. I found that quite interesting.

Even more interesting to me is how they come up with in-game percentages as to who is going to win. We all do that to some degree. Five minutes to go and your team has a ten point lead. You are probably going to win. So it is over 50%. But is it 60%? 85%? They know. At least I would say that their guess is as good as it gets.

I want to give Jay Boice and Nate Silver credit because I'm just relaying what they say is how their group comes up with those percentage win chances. I will try to do their explanation justice. So with a bit of paraphrasing, here we go,

  • You you are ahead by 10 a with five minutes to go. The question becomes - How often have teams in that same situation done that in the past?
  • They use regression analysis based on various game situations in the past. "The past" being the scores from all of the NCAA games over the past five years. 
  • It makes a difference if that team that is ahead is really the better team, so they also factor in the pre-game win probabilities. That team currently in the lead may be more lucky than good.
  • Finally, what is the current situation? It's five minutes to go. But who has the ball. Is one of the teams getting ready to shoot free throws?
  • They don't account for everything, e.g., a player has fouled out and won't be available the rest of the game. That certainly could have an impact. 
  • There probably are a number of factors that are just too much to deal with, so they don't.
Their results are pretty impressive. I haven't checked them out in real time. Its always after a game has been played. I'll have to remember to do that. Looking at them after the fact, though, their results seem pretty impressive. You can see some of their March Madness work here: 2017 March Madness Predictions

Monday, March 20, 2017

Rating the Polls

I was going to call this week's post "March Mathness" and talk a little about the NCAA tournament. Let's do that next week. Let me go ahead, though, and apologize for the title now. I'm sure I'm not the only one to use this type of play on "March Madness". That still doesn't make it right.

There is a nice website by the name of fivethirtyeight.com. It presents information regarding polls and polling data (The 538 part comes from the fact that there are 538 electors in the electoral college.) One interesting part of the website is looking at various polls (there are a lot more than I would have imagined - they rate over three hundred polling firms).

The reason I got there is because I was trying to figure out how their site, can come up with in-game information like Arizona is ahead of St. Johns 55 to 46 with 3:38 left to play, thus Arizona has an 89% chance of winning. Wow. It's clear Arizona would probably win, but how do they come up with a percent like that? Anyway, we'll look at that next week.

I got side-tracked with a section that speaks to how they rate various polls. For example the Trump/Clinton election did not come out as most had predicted. Some polls are better than others. They rate them all. For example, one of the best seems to be the ABC News/Washington Post poll. On the other had, an organization called Research 2000 is not. An overview of their methodology is at:

https://fivethirtyeight.com/features/how-fivethirtyeight-calculates-pollster-ratings/

They don't really give enough information to show exactly how they do it. That would probably be beyond me anyway. Let me tell you something they have used in the past. It is an especially cool math application since it has a square root stuck in there.

Total Error = Square Root of (Sampling Error + Temporal Error + Pollster Induced Error)

Why don't polls come out exactly right:

  1. Sampling Error:  Sampling not enough people or not getting a representative sample
  2. Temporal Error:  The farther away it time a poll is from the event; the more error
  3. Pollster Induced Error:  Seems to be kind of a catch-all category for other things that can go wrong, such as assuming a too high or too low voter turnout.
Something else interesting they talk about is the concept of "herding". The companies that do the polling want to look good. It does not look good if they've wandered too fall away from the rest of the herd. If most every other poll has candidate A having around 55% of the vote and you predict he'll have 73%, you might make an "adjustment" to your results. Or you simply chose to not publish those results in which your company seems to be way off. 

That and other factors make it pretty complicated. Polling itself is complicated and then ranking the pollster even more so. 

I hope I've done justice to what they do. If you read what they have to see on their website you can see the complexity involved.

Next week, March Madness. Don't worry it will still be going on. In fact, it is actually March and slopping over into a little bit if April Madness. 





Sunday, January 1, 2017

World Population Growth

Here is an interesting graph. (https://ourworldindata.org/world-population-growth/) It shows the world's population growth up to the present day and then someone's estimates as to what will happen in the next few decades. It probably takes a little looking at for it to make sense. I combines two graphs in one. The horizontal axis shows the passage time in years and the vertical shows growth rates. The graph also shows the total population although these numbers are just recorded on the graph rather than being recorded on the vertical axis. As line graphs go, it's a pretty busy graph.

It seems to me that math teachers could make use of the graph in pretty much any high school mathematics class.

This actually started for me with information I found in the 2017 World Almanac. It showed population estimates going much farther back in time than this graph shows. You could look at that information as a set of ordered pairs with years being represented as x-values and world population (in billions) as y-values. The almanac states that in the year one the population was an estimated 300 million. That gave me an ordered pair of (1,0.3). Proceeding in this manner gave me ordered pairs of (1,0.3), (1250,0.4), (1500,0.5), (1804,1), (1927,2), (1960, 3), (1974,4), (1987,5), (1999,6), (2011,7).

Just using the raw data, an Algebra I class might simply write ordered pairs, or without seeing the above graph, choosing appropriately labeled axes to graph the data.

Higher math classes could look at finding an equation to model the data. I had more trouble than I thought I would. I guess that is because, as the graph shows, the rate has varied over time just in the last couple centuries, let alone millennia. Leaving out the first few ordered pairs and adjusting the data such as changing (1804,1) to (0,1) and so on, I was able to find an equation that had a correlation of r = .9647. Students could maybe experiment with similar things to get a best fitting curve.

Calculus students would be able to examine the blue population growth curve and discuss how it ties into first and second derivatives. It is interesting that the person making the future projections seems to think our current point in time seems to correspond to an inflection point. Students could discuss what that really means and mathematically and socially.

Tuesday, November 8, 2016

Election Day

This just happens to be election day. Let's follow up on my posting last week about polls. Today there won't be polls, there will be projections. They change the name, but they're really the same thing. It turns out they can be wrong.
  • I read an article that said there was a primary several months ago in which pollsters took data to say that Clinton had a 99% chance of winning. Sanders ended up with a narrow victory. I guess you can't say he was wrong. Some things that are predicted to happen one percent of the time, do happen. Still probably embarrassing for those pollsters, though. 
  • In 1948, Harry Truman famously held up a newspaper declaring that "Dewey Defeats Truman". He didn't. Dewey had such a lock on it. As George Gallup Jr. said about this, "We quit polling a few weeks too soon." That'll do it.
  • In 1936, Literary Digest conducted a survey of its readership. It picked Alf Landon over Franklin Roosevelt. It turns out that Literary Digest (which has since gone out of business for obvious reasons) mostly appealed to a higher-income type person. That skewed Republican, thus predicting a President Landon.
  • In 2000 the television networks declared Al Gore the winner of Florida. That was all he needed to be president. They had to retract that, declaring the race "too close to call". Overnight the networks declared George Bush the winner. Later, back to "too close to call". 
Last time we looked at how the polls work. A newscaster might say, "Candidate X is at 57% with a 3% margin of error". They usually don't mention the level of confidence. I thought they use a 90% confidence level. I saw something lately that said it is usually 95%. Regardless, they're pretty confident. But they aren't certain. 

Let's take that example and use a 95% confidence level. 

Candidate X is at 57% with a 3% margin of error translates to:

We are 95% sure that he is somewhere between 54% and 60%.

There is actually more to it than that. Consider a bell-shaped curve peaking at 57%. Of all the possible outcomes, 57% is most likely. Then 56%, then 55%, then 54%. Even 53% or below is not out of the question. Very unlikely, but not out of the question.

So if Candidate Y is at 45%, he/she is probably going to lose. However, it won't be because the 3% margin of error says he has to.

Pretty confusing. No wonder the pollsters get embarrassed every once in a while.

Monday, October 31, 2016

Political Polls and margin of error


At the time of my writing this, it is about a week until the election. It's Trump vs. Clinton - Duel of the Century. I thought I would look into political polls as a mathematics application this week. Specifically, let's look at what is called the margin of error.

I looked at a couple of what I think are reputable websites. However, they seemed to not quite get this concept. For example, something like this was stated by a couple of sites:

A poll states that candidate A is at 52% with a margin of error of +/- 3%. This means the candidate could actually be polling anywhere from 49% to 55%.

Unless I've been lied to in my past math classes, I believe this is wrong information. This is a common misconception, but I didn't think I would find news agencies writing this.

He is what I believe is the correct scoop. Polls usually have a confidence level. Part of the confusion is when CNN, NBC, etc mention their polls, they don't talk about this. Anyway, for most polls it is 90%. So, in actuality, a much truer fact is that there is a 90% chance that that candidate A is between 49% and 55%. She (or he - I'll stick with "she" the rest of the way so I don't have to mention both genders each time. Why "she" rather than "he"? I flipped a coin. Seriously.) is probably in that range, by she can't be certain of that.

You can never be certain of polls. Common sense tells you that you can't have absolute certainty. If there are millions of voters in the U.S., and your survey covers a few thousand, how do you know you didn't just happen to survey only ones that are against candidate A. Yes, unlikely, but it could happen. So if a poll states A is ahead of B, 57% to 42% with a margin of error of 5%, it's all over, right? No, it isn't. It's not looking good for B, but it's not all over.

We see surveys during election years a lot, but we see them often at other times without knowing it. The government's unemployment reports, bestselling books, the top TV shows for the week are all done by random sampling of a relatively small sample.

Students could figure out the margin of error. It goes like this:

Margin of error = z x squareroot(p(1-p)/n). The z-value is based on how accurate you want your poll result to be. You would have to look that up. The value of p is your polling result and n is the number in your sample. (Oddly, the number in your total group, whether it is the entire U.S., the state of Oregon, or your bowling league, has nothing to do with the answer.)

Common sense tells us that there is a trade-off. The more exact you want to be, the wider your interval is going to end up being. I might be able to state, from a recent survey of adult males, that I am 90% certain the average height of all adult males is between 5'7" and 5'11". One the other hand, if I want to be 99.99% certain, I might only be able to state that the average height is between 3' and 8'. You gain in certainty and you lose in precision.

Let's try one out.

We polled 1,000 people. Of those, 560 said they would vote for Candidate A. So, she is polling at 56%. We want to be 90% certain of the range her number would actually land in. Looking up the 90%, we find a z-value of 1.645.

1.645 x squareroot(.56(1-.56)/1,000) = .0258. If we round it to 2.5%, she is 90% sure of her actual number being between 53.5% and 58.5%.

Just for fun, here are some other possibilities.

Suppose we chose a confidence level of 95%:

95% corresponds to z = 1.96, so
1.96 x squareroot(.56(1-.56)/1,000) = 3.1%, giving a range of 52.9% to 59.1%

Suppose we take our original example and assume we surveyed twice as many people:
1.645 x squareroot(.56(1-.56)/2,000) = 1.8%, giving a range of 54.2% to 57.8%

I was right. That was fun.







Monday, October 24, 2016

Standard Deviations and Baseball

Its World Series time and I feel compelled to stick with a baseball theme this week. I've considered this application since I was not much more than a child. I wasn't sure how the math on it would work, and I'm still not certain, but I thought it would be worth exploring.

Batting averages are the ratio of hits to times at bat. So getting one hit in four times up to bat gives a batting average of .250.

It would make sense that the overall batting average in baseball might vary over the years. Things have changed since it started in 1869. There used to be no night games. That is mostly because the electric light hadn't been invented yet. Night games have made it harder for hitters. Although, they've outlawed spit balls. That has made it easier.

Does it all even out? Apparently not. There used to be quite a few batters that hit .400 or better for a season. No one has done that in the past few decades, though. I've wondered if there a way to even things out mathematical. I've seen some attempts at this.

I found a person's website that has the major league batting average for each season. Over a century worth, it is at .263. The highest year ever was 1894 when it was .309. So maybe a player that year could have their batting average dropped by .046 (.309 - .263 = .046). Similar adjustments could be made for players of each year.

Not a bad idea. I've seen other similar methods. However, I've thought that some measure of variance should come into play. I've had a theory that the standard deviation of the batting average statistics have been going decreasing over the years. So, there were more .400 hitters in the past, far above the league average, but I would guess that back then there were also more hitters far below the league average.

Why might that be? Now there are scouts going to colleges, high schools, Japan, Dominican Republic, etc. looking for possible talent. In the early days, they took what they could get. It wasn't necessarily the best baseball talent. Someone might come in from the coal mines, look pretty good, and you sign him to a contract. Over the years the process has improved.

To take a shot at that proving my theory, I used a website that showed the league average for each year. I then entered twenty years worth of yearly batting averages and found the standard deviation. Its not perfect, but I think it kind of backs me up. Here we go:

1871-1900  Standard deviation = 15.91
1901-1920  Standard deviation = 10.66
1921-1940  Standard deviation = 7.38
1941-1960  Standard deviation = 3.76
1961-1980  Standard deviation = 7.91
1981-2000  Standard deviation = 5.84
2001-2012  Standard deviation = 5.15

So to really do this right, I probably should find the standard deviations of each individual year using each player, rather than using the year as a whole. However, that seemed like a lot of work, so I settle for this. I bet there is some data base that has all the averages and the capability of adjusting the mean averages and the standard deviations for each year and adjusting each player's batting average accordingly. It won't be me, but somebody should take that on.



Tuesday, March 1, 2016

Super Tuesday

Today happens to be Super Tuesday. A few days ago I saw a news release on cnn.com that I thought was interesting.

Cruz has the backing of 28% of Republican voters nationwide, unseating Trump, who won the support of 26% in the latest NBC News/Wall Street Journal poll. But Cruz's 2-point edge is within the poll's margin of error, and it's not clear if the survey captures real movement in the race or is simply an outlier. The results are a major change from last month's NBC News/Wall Street Journal poll, when Trump held a 13-point lead over Cruz, 33% to 20%.

"I have never done will in the Wall Street Journal Poll. I think somebody at Wall Street Journal doesn't like me, but I never do well in the Wall Street Journal poll," Trump said. "So I don't know. They do these small samples and I don't know exactly what it represents."

Another Wall Street journal poll released Wednesday found registered voters to be divided on whether the Senate should vote this year on a replacement for late Supreme Court Justice Antonin Scalia. Forty-three percent believe the Senate should vote on a replacement this year rather than wait for President Barak Obama's term to end, versus 42% who oppose a vote.

The NBC New/Wall Street Journal pollsters contacted 800 registered voters for the question on the Supreme Court, with a margin of error of  +/-3.5 percentage points, and 400 Republican primary voters for the Republican field, with a margin of error of 4.9 percentage points.

There are some good mathematical points to be made in this article.
  • The first paragraph could lead to a discussion of what an outlier is. 
  • The second paragraph shows that candidate Trump should probably do some research on how polls and "these small samples" work.
  • The final paragraph speaks to the concept of margin of error. Note that the margin of error was greater when fewer people were polled.
The margin of error is usually not well explained when used in news reports and is generally misunderstood. A poll result of 43% with a margin of error of 3.5% would imply a range of 39.5 to 46.5%. However, the true result isn't necessarily within that range. Pollsters tend to state that the true result would be in that range with a probability of 90%, 95%, or 99%. The article doesn't say what that probability is.

We can figure it out, though. With a sample size of n, a 99% level of confidence yields a range of 1.29/sqrt n. A 95% level is 0.98/sqrt n. A 90% level of confidence is 0.821/sqrt n.

Most polls use a 95% level of confidence and that, in fact is the case here. Note that 0.98/sqrt(800) = 3.46% and that  0.98/sqrt(400) = 4.9%, which agree with the values given in the article.

Monday, February 8, 2016

Intentional Fouls

There is a basketball statistic known as offensive efficiency (OE). It is the number of points a team score in one hundred possessions. The top teams as of right now is Golden State with and OE of 113.2. Worst is the Philadelphia 76ers at 94.5. They happen to also have the best and worst records, respectively, in the league.

In the basketball world there has been a debate about intentionally fouling a team's worst foul shooter, let him shoot two foul shots with the idea he is probably going to miss one or both of them. Is this a good strategy? Some math can be used to try to find an answer to this.

This debates centers mostly around Andre Drummond of the Detroit Pistons. he's good at some basketball things, but shooting free throws isn't one of them. He currently makes 34.9% of them.

A good probability problem might be, What is the probability makes at least one, i.e., he makes the first (event A) or second (event B) shot?

Pr(A or B) = Pr(A) + Pr(B) - Pr(A and B) = 0.349 +0.349 - (0.349)(0.349) = 0.576 = 57.6%

But I digress. Let's answer the specific question of how would Andre do just shooting free throws each time down the court? The probability of him:

Missing both = 0.651 x 0.651 = 0.424
Miss the first and make the second = 0.651 x 0.349 = 0.227
Make the first and miss the second = 0.349 x 0.651 = 0.227
Making both = 0.349 x 0.349 = 0.122

His expected value of scoring for a possession is:
0(0.424) + 1(0.227) + 1(0.227) + 2(0.122) = 0.698

Per one hundred possessions that would be 69.8 points. That compares to the Pistons team OE of 102.5. So yes, based on this, foul him. However, there are other factors to consider. Each player can only commit six fouls before being disqualified. You may have to pick your times. That is what the opposing teams have been doing.

How about the second worst free throw shooter in the league - DeAndre Jordan at 42.1%? Doing the same math, in a hundred possessions, his free throws would account for 84.2 points. His team, the Clippers, usually score 106.2, so again - foul him.

The third worst shooter is Dwight Howard. his free throw shooting would gain 109.8 points. His team's (the Houston Rockets) OE is only 104.2.

By my calculations, this fouling strategy would only be effective with two players in the entire league.

Aspiring NBA basketball players - practice your free throws.

Monday, February 1, 2016

Revolutionary War Cryptography

Spies have been around for a long time. Part of being a good spy is being able to send and receive coded messages. Lately, mathematicians have had a major role in trying to break these codes. There was a major movie, The Imitation Game, about mathematician Alan Turing and his breaking of the German Enigma code in World War II.

Codes go way before that, of course. I'm reading George Washington's Secret Six: The Spy Ring that Saved the American Revolution. The title seems a little overstated, but then again, I'm not done with the book yet. There are a couple of interesting items.

Invisible Ink - I always thought that was a made up thing. The author of the book says,

"The practice of writing with disappearing inks was nothing new. For centuries people had been communicating surreptitiously through natural and chemically manipulated inks that became visible when exposed to heat, light, or acid. A message written in onion juice, for example, dried on paper without a trace, but became readable when held to a candle. Secret correspondence in the British military often had a subtle F or A in the corner indicating to the recipient whether the paper should be exposed to fire or acid to reveal it message."

Interesting. Also, the coding itself was not as complicated as it is now, but still pretty effective. The book says that Benedict Arnold, when communicating with his British contact used,

"invisible ink and a book-based code. He based his code on two books: William Blackstone's Commentaries on the Laws of England and Nathan Bailey's An Universal Etymological English Dictionary. Each word was denoted by three numbers separated by a period. The first was the page number, the second was the line, and the third was the position of the word, starting from the left margin, in that line. for example, 172.8.7s stood for "troops": page 172, line 8, seventh word in. The s at the end simply made it plural."

The method is very clever. Although, if he was clever enough, Benedict wouldn't be so well known to us today.


Monday, January 4, 2016

Extra Points

The NFL regular season is over now. They did an experiment with extra points this year and it might be interesting to see how it turned out.

The NFL felt that extra points after touchdowns had become kind of boring. After scoring a touchdown, there is a kick from the two yard line worth one point which was pretty much always successful. During the 2013 seasons it was been successful 99.6% of the time. In 2014 it was successful 99.3% of the time. Not a sure thing, but pretty close.

Teams also had the option of going for two points by running or passing it in. That is harder to do, so, despite the lure of two points teams didn't usually opt for it unless it was near the end of the game and a team really needed those points.

To spice things up they decided to keep the two point option just as it was, but the kick had to made from the fifteen rather than the two yard line. Now that the seasons is over, we can look at how it turned out and what a team might strategically decide to do next year.

Starting from 2013 to the present season, since there were no changes, we would expect the two-points to stay about the same.

2013 - 33 for 69 - 47.8%
2014 - 28 for 59 - 47.4%
2015 - 45 for 94 - 47.8%

(I couldn't find a source that gave the stats for the entire league, so I had to add the team totals. I double checked, so I believe I'm correct on these numbers.)

While we would expect the success rate to stay about the same, it is kind of amazing to me that the percentages came out so close. We do see an increase in the number of tries last year, which is probably due to that fact that the one-point tries are not as automatic as they used to be.

One-point conversion success did go down this year, but not a lot.

2013 - 99.6%
2014 - 99.3%
2014 - 94.2%

So now, a question might be, "Should we just go for two every time now?" That boils down to what is our expected value for the various attempts.

2013 - One Point Attempts - 0.996x1 = 0.996. Two Points Attempts - 0.478x2 = 0.956.  Kick it.
2014 - One Point Attempts - 0.993x1 = 0.993. Two Points Attempts - 0.474x2 = 0.948.  Kick it.
2015 - One Point Attempts - 0.942x1 = 0.942. Two Points Attempts - 0.478x2 = 0.956.  Go for two.







Monday, December 14, 2015

Rating Baseball Players

I stumbled onto something called Elo Rater. It is a way of rating former or current baseball players if they were to face off in a head to head match up. It was developed by Arpad Elo. He was born in what was at that time Austria-Hungary in 1903 and passed away in 1992. Arpad was a physics professor at Marquette University and was an avid chess player. He developed his rating system originally to rank chess players.

Frankly, I don't completely understand every bit of this. There are original point values for the players. I'm not sure how those are determined. And I don't know who decides how it is determined that one player goes against another. It looks like maybe people can go to the website and pick a couple players and they play each other. Since he was a college professor, I'm going to assume he knew what he was doing. Also his ratings seem pretty accurate. His top five hitters of all time are:

1. Babe Ruth
2. Stan Musial
3. Ty Cobb
4. Lou Gehrig
5. Mike Schmidt

You can check out his full lists at http://www.baseball-reference.com/friv/elo.cgi

Here is an example used on the site.

RA is the rating for Player A and RB is the rating for Player B. Working out the probability that Player B wins where RA =  2450 and RB = 2500:

P(B wins) = 1 / (1 + 10^((RA - RB) / 400)) 

= 1 / (1 + 10^((-50) / 400)) 

= 1 / (1 + 10^(-0.125)) = 

= 0.571

To analyse how these come out means looking at fraction exponents and negative exponents. You can find the details of the process at http://www.baseball-reference.com/about/elo.shtml

Tuesday, November 17, 2015

Chance of Being Undefeated

They said that now, at the midpoint of the NFL season, there are still three undefeated teams. That is the most there have ever been at this point in the season. What is the chance we have at least one of those teams staying undefeated for the whole season?

There are seven games to go for each team. First we would want to know what the chance of winning a single game would be. I'm going to go with an 80% chance. That sounds about right. They're undefeated at this point, so obviously pretty good, yet they wouldn't be invincible.

The chance of any one team of winning their next seven games is (0.8)^7 = 0.21. Now what is the probability of New England, or Cincinnati, or Carolina remaining unbeaten? The easiest way to deal with a three team "or" problem is to look at it negatively. There is a 1-0.21 = 0.790 chance of a team having at least one loss the rest of the way. The chance all three are defeated is 0.493. So, there must be a 1-0.493 = 0.507 of that not being the case. So chances are at least one team will be undefeated - a 50.7% chance.

It could be done the straight forward way, but its a little more cumbersome.

0.21+0.21+0.21-(0.21)(0.21)-(0.21)(0.21)-(0.21)(0.21)(0.21)+(0.21)(0.21)(0.21) = 0.507 = 50.7%

For full disclosure, since I thought about it yesterday, Cincinnati lost on Monday Night Football. I was just giving them the win. I probably shouldn't have done that.

Now we're down to two teams. There is still a (0.8)^7 = 0.21 chance of any one team winning seven in a row. The chance that either Carolina or New England does it is:

(0.21)+(0.21)-(0.21)(0.21) = 0.347 = 34.7%

If you take exception to my 80% single win assumption, it's easy to substitute another value. This makes a bigger difference in the final probabilities than I would have thought. For example, if you figure the chance of a team of this caliber wins a single game is 85%, the chance one of the remains unbeaten is now 53.9%. If 90%, the chance of having an unbeaten is all the way up to 73.8% .

Sunday, November 1, 2015

The Law of Large Numbers and MVP's

As you do repeated trials, the mean average of those trials will approach the theoretical mean. Or if you are looking at a sample, as your sample grows, its mean average will get closer to the mean of the entire population.

The law of large numbers is a good title for this. Many have the idea that the law of probability would imply that a .250 hitter that goes 0 for 3 is now "due" for a hit. Flipping a coin 10 times means you'll get 5 heads and 5 tails. For most of us, our own life experience would show statements like these to be incorrect. It would not be weird for a coin being flipped 10 times to have 3 heads and 7 tails. However, we would think something was up if we flipped it 1,000 times and got 300 heads and 700 tails.

Like many math teachers, I would have classes do some coin flipping experiments. Always a fun day. For years I would write down the results and keep a running total. I don't know where that is now. I wish I had kept that. I was up to something like 20,000 flips. It wasn't 50-50, but pretty close. Maybe something like 49.7% to 50.3%.

I'm reminded of this in something I read in "The Signal and the Noise: Why So Many Predictions Fail - But Some Don't" by Nate Silver. It's an interesting book. One example: At what is a major league baseball player at his peak? You can make a pretty good case for it being 27 years of age. To make his case, he looked at 50 MVP award winners. Granted, 50 is possibly not to be considered a "large number", but it's what was used in this case.

"A baseball player...peaks at age twenty-seven. Of the fifty MVP winners between 1985 and 2009, 60 percent were between the ages of twenty-five and twenty-nine, and 20 percent were aged twenty-seven exactly."

This doesn't prove anything for sure, but then again, surveys never do. I would think we would have a better idea as we can look at additional MVP's in coming decades, giving more applicability to the law of large numbers. Also, the results might have been more convincing if they didn't include Barry Bonds steroid-assisted MVP awards in his late thirties.


Monday, October 12, 2015

The Sophomore Jinx

The sophomore jinx, or sophomore slump, takes place when the second round is not as good as the first. The second album, the second season, doesn't seem to be quite as good. They got everyone's hopes up after a great debut. What's with that?

Does it really happen, anyway? Maybe its just hit or miss. The Grammy Award for Best New Artist in 1964 was a group called the Beatles. Good call. But the year before, the Best New Artist was Ward Swingle. First let's look at the case for there being such a thing as a sophomore jinx. With a quick look at the internet one can find plenty of examples that seem to support this idea.

Album sales by some pretty well known names:

Terence Trent D'arby - Album Number One - 12 million, Album Number Two - 2 million.
Spin Doctors - Album Number One - 5 million, Album Number Two - 1 million.
Christopher Cross -  Album Number One - 5 million, Album Number Two - 500,000 thousand.
Hootie and the Blowfish - Album Number One - 16 million, Album Number Two - 3 million.

You get the idea. Aaron Gleeman in an article titled The Sophomore Slump looked at all of the Rookie of the Year award winners, comparing their first and second seasons by using a baseball statistic called win shares. He found that 73 of the winners got worse in season two, while only 37 improved. Four stayed the same.

Rick Sutcliffe was the National League Rookie of the Year in 1979. Overall, he had a fine career, winning 179 games. His first year he won 17 games and lost 10. He gave up about three and a half runs a game. Next year he won 3 and lost 9 and gave up about five and a half runs a game.

This so-called sophomore jinx, can be explained at least in part statistically with the concept of the regression to the mean. The dictionary says, "In statistics, regression toward (or to) the mean is the phenomenon that if a variable is extreme on its first measurement, it will tend to be closer to the average on its second measurement.

If we flip a coin 100 times and get 63 heads, would we do better next time? Yes, maybe. But probably not. But if we got 40 heads on a first try, chances are, next time we'll see an increase. In either case we go back toward or regress toward the mean.

The examples we have seen have something in common. All of these first years were very good. They got our attention. People are wondering what they'll do for a follow up. All of these burst upon the scene with a great debut. Perhaps they were far above what their usual production would be. That can happen, but chances are in any given effort, we will do what our historical average would suggest.

Are there cases where the sophomore slump doesn't happen? Consider the baseball player that has a first year that is a bit below what he is capable of. Likely, he will improve the next year. The public really didn't notice his first year because it was nothing spectacular. We we're all paying attention to the Rookie of the Year winners. 

Friday, September 18, 2015

Ryan Braun and Steroids

Ryan Braun is an outfielder for the Milwaukee Brewers. He won the MVP award in 2011. At the end of that season he was accused by Major League Baseball of taking steroids. However, he got off on a technicality. He was accused again in 2013. This one stuck and he was suspended for the rest of the season. That time he admitted it and took his punishment. He has played almost two full seasons since his suspension.

What is interesting in Braun's case is that he is still in the prime of his career. He is 31 years old. Many that have been accused of taking steroids were near the end of their career. If they did come back from a suspension, a decrease in their statistics could be because they are no longer using steroids or just because of father time. A decrease in Braun's statistics would seemingly be due only to him now playing clean.

To compare his statistics pre and post-suspension would be an interesting exercise. It wouldn't be helpful to look at the totals as he played almost seven seasons before the suspension and only two seasons after. However, you could translate those time periods into single 600 at-bat seasons. That is what I did. You can do so by looking at the grand totals and setting up proportions. Also helpful in this exercise is knowing that the definition of batting average is the number of hits divided by the at-bats.

Algebra students would have plenty of opportunity here to review proportions. Let me just skip the messy stuff and go right to the final stats.

Pre-Suspension Statistics:
At-bats 600, Runs 104, Hits 187, Doubles 38, Home Runs 34, RBIs 110, Batting Average .312

Post-Suspension Statistics:
At-bats 600, Runs 90, Hits 166, Doubles 33, Home Runs 26, RBIs 96, Batting Average .277

You could then ask students what they make of these statistics. Some, perhaps with some leading by you, might suggest looking at the percentage decrease. This turns out to be quite interesting. You can easily make the claim that a player is 86 to 87% as effective without using steroids. At least that seems to be the case with Braun in pretty much every area. I've compare post to pre-suspension statistics and changed them into percentages. Check this out:

Runs 87%, Hits 89%, Doubles 87%, Home Runs 76%, RBIs 87%, Batting Average 89%. Its kind of surprising these numbers are all in the same ballpark, so to speak.

Anyway this might make a good review of proportions, takes a topic they've all heard about, and gets students to do some statistical analysis.

Monday, September 14, 2015

Smartest Presidents

I saw an article online which listed the IQ's of each of our presidents. By their own admission, they were doing a bit of guesswork. Since IQ tests weren't developed until about the time of our 26th president, Teddy Roosevelt, there isn't a lot of hard data that can be used. I would think we would take these numbers with a grain of salt. I have some disagreements with a few of these placements and you probably do, too. In spite of that, here we go:

The Top 5 and their estimated IQ scores:

1. John Quincy Adams - 168.8
2. Thomas Jefferson - 153.8
3. John Kennedy - 150.65
4. Bill Clinton - 148.8
5. Woodrow Wilson - 145.1

And the bottom 5:

Andrew Johnson - 125.7
George W. Bush - 124.9
Warren G. Harding - 124.3
James Monroe - 124.1
Ulysses S. Grant - 120.0

We might well ask just how smart these guys really are. Since the mean IQ score is taken to be 100, just like Lake Wobegon, they are all above average. So, President Grant was above average. But was he just a little above or way above?

We could get an idea from looking at how many standard deviations away from the mean he is. Taking the standard deviation for IQ scores to be 16, we see that Grant is 20/16 = 1.25 standard deviations above the mean. Consulting a Z-score table, that puts him in the 89.4 percentile. Pretty good. Even if Ulysses wasn't the sharpest guy ever, it must take a certain amount of intelligence to get elected president twice and to win a war.

What about John Quincy? J.Q. is literally off the chart. Let's go with the runner-up Thomas Jefferson. He is 53.8/16 = 3.36 standard deviations from the mean. This puts him in the upper 99.96 percentile. He's smart. Not John Quincy Adams smart, but smart.

The complete list can be found at http://us-presidents.insidegov.com/stories