Showing posts with label standard deviation. Show all posts
Showing posts with label standard deviation. Show all posts

Monday, 26 July 2010

Standard Deviations of Sums of Distributions

A week or so ago I was at a textbook selection conference with a couple of really good teachers, and one of them (thanks, Dru) pulled out a copy of Robert Hayden's, "Advice to Mathematics Teachers on Evaluating Statistics Textbooks." I mention it now because it has two good pieces of advice. (Ok, it has way more pieces of good advice than that, but I'm mentioning these two in particular)

The first, I hope I follow, "...make sure the textbook mentions assumptions and teaches students to check them rather than make them." One of the ways I try to get students to check assumptions is to make them understand, as much as possible in the limited time of a AP course, the WHY. In order to do that, I frequently violate one of Professor Hayden's other pieces of wisdom; "Be wary of an author who is not familiar with enough real data sets to illustrate a textbook."

Ok, I'm gonna claim some "weasel" room here. First, I'm thinking more along the lines of an exercise to help the students understand why checking independence is so important, and not writing a textbook. Second, even Professor Bob himself says "While there may be places (such as Anscombe’s regression examples , in which a skillfully fabricated batch of numbers illustrates a pedagogical point,.." Ok, so the "skillfully" may not apply to what follows, but I hope the fabricated data at least help drive home a "pedagogical point".

I begin with two simple data populations, X= {1,1,1,2,2,2,3,3,3} and Y= {1,1,1,3,3,3,5,5,5}. Students who have learned the "Standard Deviation as Distance" approach can quickly check and find the standard deviation of the X population (or using a calculator) is sqrt(2/3)or appx .8165. For Y the std. dev. is 1.633. Perhaps for what we will be doing, we remind them that the variance of each is the square of the standard deviation, so Var(X)=2/3 and Var(Y)= 8/3.

So what happens if we add or subtract the populations? It all depends! If the populations are independent, then any X and any Y may (must?) be associated with equal probability. I illustrate this by pairing one of each X value with one of each Y.. (is it possible to have two distributions be independent without this type of each x with each y association?)

X___1___1___1___2___2___2___3___3___3

Y___1___3___5___1___3___5___1___3___5.

and the sum and differences are then

X+Y =2___4___6___3___5___7___4___6___8 and

X-Y =0__-2__-4___1__-1__-3___2___0__-2

I think it is worth drawing the two resulting distributions because many students will NOT see that these are distributions are reflections of each other. So they should have exactly the same standard deviations (this takes a moments reflection for some students).




Wow, that's good news. If the populations items are independent of each other in the way they are combined, it doesn't matter if you add them or subtract them, the spread is the same since the two distributions are symmetric, which means the standard deviations should (and are) the same, about 1.8257. Even better, we can point out that the variance, 10/3, is simply the sum of the original variances, 2/3 + 8/3. For me it is worth pointing out this "Pythagorean" relationship, [StDev(X+Y)]2=[StDev(X)]2+[StDev(Y)]2, IFF X and Y are independently associated.....(oops, I have been called out on this mistake... The statement is true IF x and y are independent, but also in any situation in which the correlation coefficient is zero... which does not necessarily require independence...see comment from "gasstationwithoutpumps" below... "mia culpa" and thanks to "gas..."

BUT... what if the original populations were NOT independent. (quick, think of two data sets that you would really combine in real life that are totally independent...better yet, send your ideas in the comments)

Well they might have a positive or a negative correlation, so we slightly rearrange our data sets and group lower numbers somewhat together (no Ones with the fives) like this..

X___1___1___1___2___2___2___3___3___3

Y___1___3___1___1___3___5___5___3___5.

Now our sums and differences are

X+Y =2___3___2___3___4___5___6___5___6 and

X-Y =0__-1___0___1___0__-1___0___1___0

We recognize quickly that the sets no longer have the same shapes. The distribution of sums is almost uniform with the peaks at the ends, while the difference distribution has two peaks closer to the center .

So what are the spread measures now. The standard deviation of the summation distribution is 2.26 or the square root of the variance of 46/9. The differences have a standard deviation of 1.247, the square root of a variance of 14/9, a really big difference. In fact, we help the students notice that the variances are the same distance from the equal variance of 10/3 = 30/9 when the populations were combined independently. The distribution of sums variance is 16/9 higher, the differences are 16/9 lower. Is this just a curious coincidence...(by now my students know that almost NOTHING I bring up is a "curious coincidence" ).

So how can we explain this difference. Slowly you lead their thinking...."If the distributions are NOT independent, they must be dependent,.... and there must be some relationship,..... some measure of how UN-independent they are." Eventually they will think of the correlation coefficient, r. In this association between X and Y they have a positive correlation of 2/3 ... can that help. If the relationship when the association was independent is "Pythagorean", maybe we can look for some extension of the Pythagorean theorem to help... Can we find something like the Law of Cosines that would tie the package together? After all, we need something that will add 16/9 to the sum distribution, and subtract the same amount for the differences... I can't imagine that I would have kids who would see this, and will probably lead them to observe that StDev(X+Y)=[StDev(X)]2+[StDev(Y)]2+2 r [StDev(X)][StDev(Y)]. They can quickly test that the change of sign leads to
StDev(X-Y)=[StDev(X)]2+[StDev(Y)]2 - 2 r [StDev(X)][StDev(Y)].

I hope before I get to this point I have laid a foundation for this by giving a short presentation based on a blog from John D Cook at "The Endeavor" that shows this geometrical relation between the correlation coefficient and the cosine of an angle. I hope to write a blog about this relationship in a more vector sense later.

All of this follows in the wake of a warning about non-real data from Professor Hayden, so it is important to follow up with real data that should bare this out. I'm thinking something simple like their own age in months and height. If it is true for all data sets, it should be true with the measures we have about them; but I am very willing to consider suggestions about a more appropriate data base.

Monday, 12 July 2010

Standard Deviation as Distance

Early today I had a conversation with another HS stats teacher that reminded me that when I was writing about vectors a while back I had not covered two nice uses in Stats. I hope to correct one of those today.

As we were talking I bemoaned the fact that few introductory textbooks seem to really help kids to develop any intuitive idea of what the standard deviation is or how it works. As we talked, I mentioned that I thought there was a geometric approach to the standard deviation that might help make it more clear. You be the judge.

I think the standard deviation is most easily approached as a distance (more specifically a sort of average of distances). Most high school stats students can quickly find the distance between two points on the plane using the square root of the sum of the squares of the differences (deviations) in each direction (dimension). For those who have never been introduced to it, only a few moments convinces them that it can generalize to n-dimensions. And in a few short minutes they can be finding the "distance" between (point)vectors in any number of dimensions, and many can quickly invent a shortcut to the calculation using the list functions of their calculators.

So why does the standard deviation as a distance make sense? The standard deviation is a measure of how much the data items "disagree" with each other. Start with two measures, and for the moment we use the unconventional notation of calling one of them x1 and the other y1. Now if they agree perfectly, then they lie on the line y=x. If they don't, then they will be off the line by some distance. We begin by finding that distance. The perpendicular from the line y=x to the point (x1 ,y1) would cross y=x at the point where the x and y values were the average of x1 and y1, or at a point we call (xbar,xbar). That means the distance of the point (x1 ,y1) from the line y=x is just

Now if all our data sets had only two values (and statistics was REALLY EASY) then we could use this "distance" measure as a "standard measure". But one of the funny things about distance is that it grows with dimension, "sort of"... here is what I mean. In one dimension, the distance from (0) to (1) is one unit. In two dimensions the distance from (0,0) to (1,1) is farther, it's the square root of two. In three dimensions the distance from (0,0,0) to a point one away in each dimension is the square root of three. This would meant that the data set {1,3} would seem to be "less spread out" than {1,1,3,3}, which seems like a bad thing. To compensate, we simply divide this Pythagorean distance result by the square root of the dimension.

In effect then, the standard deviation of a population of values is the distance between the n dimensional points A={x1,x2,x3..xn) and B= (x-bar,x-bar,.... x-bar) divided by the square root of n. In truth, it would seem there was no need to memorize a formula when the student understands it as a "mean distance".

As a happy coincidence, John Cook at The Endeavour web site just posted a blog about the relationship between vector geometry and statistics when finding the standard deviation of a sum or difference of two distributions. A must read for intro stats teachers who want to be able to explain what happens (and why?) when the distributions are NOT independent.

Monday, 6 October 2008

Standard Deviation Computation in One Pass


One more new (to me) item in my explorations of the standard deviation. I came across a one-pass method of computing the sample standard deviation which seems to be exact, not an approximation.

Normally the way we compute the sample standard deviation requires two passes through the data, one where we calculate the mean, then again to find the deviations from the mean to calculate the RMS of the deviations. This method is good for most cases where the data comes to you in a set and complete. But in cases where you have data streaming in constantly, say as from the Hubble telescope or something, having to go back and recalculate the total each time may be time consuming. What is needed is a way to take the old results and produce the new result with the one piece of data added.

I found just such a result at a page From John D Cook He writes, "This better way of computing variance goes back to a 1962 paper by B. P. Welford and is presented in Donald Knuth's Art of Computer Programming Vol 2, page 232, 3rd edition. Although this solution has been known for decades, not enough people know about it. ...It is not obvious that the method is correct even in exact arithmetic. It's even less obvious that the method has superior numerical properties, but it does. The algorithm is as follows.

Initialize M1 = x1 and S1 = 0.

For subsequent x's, use the recurrence formulas

Mk = Mk-1+ (xk - Mk-1)/k

Sk = Sk-1 + (xk - Mk-1)*(xk - Mk).

For 2 ≤ k ≤ n, σ2 = Sk/(n - 1)."
although I think the sigma-squared should be s2 from the values I get and the division by n-1. For those who use the Ti 84 and such calculator, this is an easy program to write, or you may download a program I quickly wrote here. It asks for each x in turn, and reports the running mean and sk after each entry.

Tuesday, 16 September 2008

A Graphic Method of Calculating the Standard Deviation



I’ve always been interested in the methods of graphic solutions to equations. Fortunately, over the years Dave Renfro has been willing to keep me in mind when he runs across old journal articles and send me copies, dozens over the last few years on just this topic. So, when I found a patent for a mechanical machine that would do standard deviations, I thought there had to be a graphic method also. I set about trying to construct one, and what follows is one such method. I have confined myself to three points, but the method would generalize to any number of data values.

The first step is to plot the points on the y-axis of a coordinate grid, and draw a horizontal line through the mean of the values. The first image shows such a construction starting with the values a=7, b=2, and c=3, with the mean, M, at 4. [If images are not sharp, click on them to open a full sized image]

Clever observers will have noticed another point on the graph at (-1,4), one to the left of the mean point, M. We will use this point to square the length of the deviation of A (the distance from M to A). To Do this we make MA the mean proportional between 1 and MA2. This can be easily done by the method of similar triangles. We make a line from the point one left of M to point A to complete a right triangle. Then if we construct a perpendicular to this line at A, it will intersect the line through the mean at a new point I call AA. The distance M to AA is the square of the length MA.. we have found the square of the deviation of point A.

Now we just need to transpose this line to another graph for safekeeping where we will add it to the other squared deviations. We repeat the previous process with points B and C, and copy each of the lengths sequentially onto a line, thus finding the sum of the squares of the deviations.


There seems to be one image of the sum of the squares of the deviations added together that doesn't come up, so you can find it here




Now we employ similar triangles again to find the mean of the sum of the squares of the deviations. We need to divide the length of the sum (Origin to CC) by n=3. To do this we drop a perpendicular from one end of the line (0-CC) that is three units long. We connect the far end (CC) to the other end of the three unit line(see figure below) and then construct a parallel to O-CC one unit away from the end. By the properties of similar triangles, we know that P-MSS is the Mean of the sum of the squares (O-CC divided by three)… we have found the variance.



Now all that is left is to take the square root, and we do this by reversing the method we used to find the square of the deviations, we find the mean proportional between P-MSS and one. So we construct a point one to the left of the segment )-MSS and then bisect this new segment to find the center. Now we create a circle with a centered at this midpoint and using the MSS + 1 as a diameter. The positive distance from where this point crosses the x-axis to the point P is the standard deviation of the three points… Which I have calculated out with two decimal points accuracy…(although I used a negative sign …Oops).. .



The method would be the same, of course for any number of points.

Saturday, 13 September 2008

A Standard Deviation Calculating Machine


Searching through some old patents looking for something else a few days ago and came across a patent for a machine that calculates the standard deviaton of a set of data (it suggested no more than 25 items). Find more at My Mathwords Page