Introduction to Statistics: Difference between revisions
| Line 97: | Line 97: | ||
It's usually possible to represent the outcome of experiments in terms of a representation involving integers or real numbers. For instance, in the case of flipping a coin, it becomes a little cumbersome to record the outcomes of multiple experiments. Let's say we flip a coin four times, and repeat this experiment twice. We might have gotten the following outcomes from the two experiments.<br> | It's usually possible to represent the outcome of experiments in terms of a representation involving integers or real numbers. For instance, in the case of flipping a coin, it becomes a little cumbersome to record the outcomes of multiple experiments. Let's say we flip a coin four times, and repeat this experiment twice. We might have gotten the following outcomes from the two experiments.<br> | ||
<math>\{HHTH\}</math> and <math>\{TTHH\}</math><br> | <math>\{HHTH\}</math> and <math>\{TTHH\}</math><br> | ||
But if we're only really interested in the number of times that the coin turns up heads in four flips, we can assign an integer value of 0 to Tails and 1 to Heads. Adding the the values together for each experiment gives us what we were after in the first place. | |||
Revision as of 20:25, 29 June 2005
Preamble
Probability and statistics is the most boring thing imaginable. It's dull, unlike fancy math -- if you're ever at a cocktail party and want to impress someone, it's a lot easier to say you're a mathematician than to say you're a statistician. It's also pretty abstract. Most thorough explorations of the discipline are in very dry books that are short on examples. If fact, if you're here it's probably because you're taking a stats class somewhere and your book and lecturer totally suck.
This aside, statistics is a central topic in modern life, and understanding just a little bit about it will help you a lot regardless of what you do or end up doing. I hope this web site doesn't suck. If it does, please leave comments on the discussion tab above.
Experiments, Outcomes and Events
The easiest way to think of probability is in terms of experiments and their potential outcomes. Many examples can be drawn from everyday experience: On the drive home from work, you can encounter a flat tire, or have an uneventful drive; the outcome of an election can include either a win by candidate A, B, or C, or a runoff.
Definition: The entire collection of possible outcomes from an experiment is termed the sample space, indicated as <math>\Omega</math>
The simplest (albeit uninteresting) example would be an experiment with only one possible outcome, say <math>A</math>. If we remember our set theory from elementary school, we can express the sample space as follows:
<math>\Omega = \{ A \} </math>
A more interesting example is the result of rolling a six sided dice. The sample space for this experiment is:
<math>\Omega = \{ 1,2,3,4,5,6 \}</math>
We may be interested in events in an experiment.
Definition: An event is some subset of outcomes from the sample space
In the dice example, events of interest might include
a) the outcome is an even number
b) the outcome is less than three
These events can be expressed in terms of the possible outcomes from the experiment:
a) : <math> \{2,4,6\} </math> b) : <math> \{ 1,2 \}</math>
We can borrow definitions from set theory to express events in terms of outcomes. Here is a refresher of some terminology, and some new terms that will be important later:
<math> \cup </math> represents the Union of two events
[[Image:[1]]]
<math> \cap </math> represents the Interection of two events
[[Image:[2]]]
<math>\{\cdots\}^{c}</math> represents the complement of an event. For instance, "the outcome is an even number" is the complement of "the outcome is an odd number" in the dice example.
<math> A \backslash B </math> represents difference, that is, <math>A</math> but not <math>B</math>. For example, we may be interested in the event of drawing the queen of spades from a deck of cards. This can be expressed as the event of drawing a queen, but not drawing a queen of hearts, diamonds or clubs.
<math>\varnothing</math> or <math>\{\}</math> represent an impossible event
<math>\Omega</math> represents a certain event
<math>A</math> and <math>B</math> are called disjoint events if <math>A\cap B = \varnothing</math>
Probability
Now that we know what events are, we should think a bit about a way to express the likelihood of an event occuring. The classical definition of probability comes from the following. If we can perform our experiment over and over in a way that is repeatable, we can count the number of times that the experiment gives rise to event <math>A</math>. We also keep track of the number of times that we perform the same experiment. If we repeat the experiment a large enough number of times, we can express the probability of event <math>A</math> as follows:
<math>P(A) = \frac{N_{A}}{N} </math>
where <math>N_{A}</math> is the number of times event <math>A</math> occurred, and <math>N</math> is the number of times the experiment was repeated. As <math>N</math> approaches infinity, the fraction above approaches the true probability of the event <math>A</math>. The value of <math>P(A)</math> is clearly between 0 and 1. If our event is the certain event <math>\Omega</math>, then for each time we perform the experiment, the event <math>\Omega</math> is observed; <math>N_{\Omega} = N</math> and <math>P(\Omega)=1</math>. If our event is the impossible event <math>\varnothing</math>, we know <math>N_{\varnothing}=0</math> and <math> P(\varnothing) = 0</math>.
If <math>A</math> and <math>B</math> are disjoint events, then whenever event <math>A</math> is observed, then it is impossible for event <math>B</math> to be observed simultaneously. Then
<math>N(A\cup B) = N(A) + N(B)</math>
Given our definition of probability, we can arrive at the following:
<math>P(A\cup B) = P(A) + P(B)</math>
At this point it's worth remembering that all events are not disjoint events. I was originally confused by events and outcomes, and this was the source of many misunderstandings. For events that are not disjoint, we end up with the following probability definition.
<math> P(A\cup B) = P(A) + P(B) - P(A\cap B)</math>
How can we see this from example? Well, let's consider drawing from a deck of cards. I'll define two events: "drawing a Queen", and "drawing a Spade". We can tell off the bat that these are not disjoint events, because you can draw a queen that is also a spade. There are four queens in the deck, so if we perform the experiment of drawing a card, putting it back in the deck and shuffling (what statisticians refer to as sampling with replacement, we will end up with a probability of <math>\frac{1}{13}</math> for a queen draw. By the same argument, we obtain a probability for drawing a spade as <math>\frac{1}{4}</math>. The expression <math>P(A\cup B)</math> here can be translated as "the chance of drawing a queen or a spade". If we assume naively (as I used to) that for this case <math>P(A\cup B) = P(A) + P(B)</math>, we can simply add our probabilities together for "the chance of drawing a queen or a spade" as <math>\frac{1}{13}+\frac{1}{4}</math>. If we were to gather some data experimentally, we would find that our results would differ from the prediction -- the probability observed would be slightly less than <math>\frac{1}{13}+\frac{1}{4}</math>. Why? Because we're counting the queen of spades twice in our expression, once as a spade, and again as a queen. We need to count it only once, as it can only be drawn with probability of <math>\frac{1}{52}</math>. Still confused?
Proof: If <math>A</math> and <math>B</math> are not disjoint, we have to avoid the double counting problem by exactly specifying their union.
<math>A \cup B = A \cup (B \backslash A) </math> so
<math>P(A \cup B) = P(A \cup (B \backslash A)) </math>
<math>A</math> and <math>B \backslash A</math> are disjoint sets. We can then use the definition of disjoint events from above to express our desired result:
<math>P(A \cup B) = P(A) + P(B \backslash A)</math>
We also know that
<math>P(B \backslash A) = P(B) - P(B\cap A)</math>
so
<math> P(A \cup B) = P(A) + P(B) - P(B\cap A)</math>
Whew! Our first proof. I hope that wasn't too too dry.
Conditional Probability
Many events are conditional on the occurance of other events. Sometimes this coupling is weak. One event may become more or less probable depending on our knowledge that another event has occured. For instance, the probability that your friends and relatives will call asking for money is likely to be higher if you win the lottery. In my case, I don't think this probability would change.
Let's get formal for a second and remember our original definition of probability.
<math>P(A) = \frac{N_{A}}{N}</math>
Consider an additional event <math>B</math>, and a situation where we are only interested in the probability of the occurance of <math>A</math> when <math>B</math> occurs. A way at this probability is to perform a set of experiments (trials) and only record our results when the event <math>B</math> occurs. In other words
<math> \frac{N_{A\cap B}}{N_{B}} </math>
We can divide through on top and bottom by <math>N</math> the total number of trials to get <math>P(A\cap B)/P(B)</math>. We define this as 'conditional probability':
<math>P(A|B) = \frac{P(A\cap B)}{P(B)}</math>
which when spoken, takes the sound "probability of <math>A</math> given <math>B</math>."
A Totally Confusing Problem that Shows how Difficult it is to Conquer Intuition
Bayes' Law
An important theorem in statistics is Bayes' Law, which states that
<math> P(B|A) = \frac{P(A|B)P(A)}{P(B)}</math>
It is easy to prove. We start with identical expressions for <math>P(A\cap B)</math>.
<math> P(A\cap B) = P(B\cap A) </math>
<math> P(A\cap B) = \frac{P(A|B)}{P(B)}</math>
<math> P(B\cap A) = \frac{P(B|A)}{P(A)}</math>
Simple rearrangement gives us the first expression above.
Independence
Two events <math>A</math> and <math>B</math> are called independent if
<math>P(A\cap B) = P(A)P(B)</math>
Random Variables
It's usually possible to represent the outcome of experiments in terms of a representation involving integers or real numbers. For instance, in the case of flipping a coin, it becomes a little cumbersome to record the outcomes of multiple experiments. Let's say we flip a coin four times, and repeat this experiment twice. We might have gotten the following outcomes from the two experiments.
<math>\{HHTH\}</math> and <math>\{TTHH\}</math>
But if we're only really interested in the number of times that the coin turns up heads in four flips, we can assign an integer value of 0 to Tails and 1 to Heads. Adding the the values together for each experiment gives us what we were after in the first place.