<?xml version="1.0"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
	<id>https://ideawaza.com/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=157.193.0.0%2F16</id>
	<title>IdeaWazaWiki - User contributions [en]</title>
	<link rel="self" type="application/atom+xml" href="https://ideawaza.com/api.php?action=feedcontributions&amp;feedformat=atom&amp;user=157.193.0.0%2F16"/>
	<link rel="alternate" type="text/html" href="https://ideawaza.com/wiki/Special:Contributions/157.193.0.0/16"/>
	<updated>2026-09-30T13:18:52Z</updated>
	<subtitle>User contributions</subtitle>
	<generator>MediaWiki 1.46.0</generator>
	<entry>
		<id>https://ideawaza.com/index.php?title=Terminal_restriction_fragment_length_polymorphism&amp;diff=41669</id>
		<title>Terminal restriction fragment length polymorphism</title>
		<link rel="alternate" type="text/html" href="https://ideawaza.com/index.php?title=Terminal_restriction_fragment_length_polymorphism&amp;diff=41669"/>
		<updated>2008-06-24T15:55:45Z</updated>

		<summary type="html">&lt;p&gt;157.193.8.204: /* Patten comparison */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;&#039;&#039;&#039;Terminal Restriction Fragment Length Polymorphism&#039;&#039;&#039; (TRFLP or sometimes T-RFLP) is a [[molecular biology]] technique for profiling of microbial communities based on the position of a [[restriction site]] closest to a labeled end of an amplified gene.The method is based on excising a mixture of [[PCR]] amplified variants of a single gene using one or more [[restriction enzyme|restriction enzymes]] and detecting the size of each of the individual resulting terminal fragments using a [[DNA sequencer]].The result is a graph image where the X axis represents the sizes of the fragment and the Y axis represents their fluorescence intensity.&lt;br /&gt;
&lt;br /&gt;
==Background==&lt;br /&gt;
TRFLP is one of several molecular methods aimed to generate a fingerprint of an unknown microbial community. Other similar methods include [[TGGE|DGGE, TGGE]], [[ARISA]], [[ARDRA]], etc.&lt;br /&gt;
These relatively high throughput methods were developed in order to reduce the cost and effort in analyzing microbial communities using a [[clone library]]. The method was first described by Liu and colleagues in 1997&amp;lt;ref name=basic&amp;gt; Liu, W, Marsh, T, Cheng, H, &amp;amp; Forney, L (1997) Characterization of microbial diversity by determining terminal restriction fragment length polymorphisms of genes encoding 16S rRNA. Appl. Environ. Microbiol. 63: 4516-4522&amp;lt;/ref&amp;gt; which employed the amplification of the [[16s|16S rDNA]] target gene from the DNA of several isolated bacteria as well as environmental samples.&lt;br /&gt;
Since then the method has been applied for the use of other marker genes such as the functional marker gene gene pmoA to analyze methanotrophic communities.&lt;br /&gt;
&lt;br /&gt;
==Method==&lt;br /&gt;
Like most other community analysis methods, TRFLP is also based on PCR amplification of a target gene.In the case of TRFLP, the amplification is performed with one or both the primers having their 5’ end labeled with a fluorescent molecule. In case both primers are labeled different [[Fluorophore|fluorescent dyes]]are required. While several common fluorescent dyes can be used for the purpose of tagging such as: 6-FAM, ROX, TAMARA, and HEX, the most widely used dye is 6-FAM. The mixture of [[amplicons]] is then subjected to a restriction reaction, normally using a four-cutter [[restriction enzyme]]. Following the restriction reaction, the mixture of fragments is separated using either capillary or polyacrylamide electrophoresis in a DNA sequencer and the sizes of the different terminal fragments are determined by the [[fluorescence]] detector. Because the excised mixture of amplicons is analyzed in a sequencer, only the terminal fragments (i.e the labeled end or ends of the amplicon) are read while all other fragments are ignored. Thus, T-RFLP is differed from ARDRA and [[RFLP]] in which all restriction fragments are visualized. In addition to these steps the TRFLP protocol often includes a cleanup of the PCR products prior to the restriction and in case a capillary electrophoresis is used a desalting stage is also preformed prior to running the sample.&lt;br /&gt;
&lt;br /&gt;
==Data format and artifacts==&lt;br /&gt;
The result of a T-RFLP profiling is a graph called [[electropherogram]] which is an intensity plot representation of an [[electrophoresis]] experiment (gel or capillary). In an electropherogram the X-axis marks the sizes of the fragments while the Y-axis marks the fluorescence intensity of each fragment. Thus, what appears on an electrophoresis gel as a band appears as a peak on the electropherogram whose integral is its total fluorescence. In a T–RFLP profile each peak assumingly corresponds to one genetic variant in the original sample while its height or area corresponds to its relative abundance in the specific community. Both assumptions listed above, however, are not always met. Often, several different bacteria in a population might give a single peak on the electropherogram due to the presence of a restriction site for the particular restriction enzyme used in the experiment at the same position.To overcome this problem and to increase the resolving power of this technique a single sample can be digested in parallel by several enzymes (often three) resulting in three T-RFLP profiles per sample each resolving some variants while missing others. Another modification which is sometimes used is to fluorescently label the reverse primer as well using a different dye, again resulting in two parallel profiles per sample each resolving different number of variants.&lt;br /&gt;
&lt;br /&gt;
In addition to convergence of two distinct genetic variants into a single peak artifacts might also appear, mainly in the form of false peaks. False peaks are generally of two types: background “noises” and “pseudo” TRFs &amp;lt;ref name=pseudo&amp;gt; Egert, M, &amp;amp; Friedrich, MW (2003) Formation of Pseudo-Terminal Restriction Fragments, a PCR-Related Bias Affecting Terminal Restriction Fragment Length Polymorphism Analysis of Microbial Community Structure. Appl. Environ. Microbiol. 69: 2555-2562 &amp;lt;/ref&amp;gt;. Background (noise) peaks are peaks resulting from the sensitivity of the detector in use. These peaks are often small in their intensity and usually form a problem in case the total intensity of the profile is low (i.e. low concentration of DNA). Because these peaks result from background noise they are normally irreproducible in replicate profiles, thus the problem can be tackled by producing a consensus profile from several replicas or by eliminating peaks below a certain threshold. Several other computational techniques were also introduced in order to deal with this problem &amp;lt;ref name=dunbar&amp;gt; Dunbar, J, Ticknor, LO, &amp;amp; Kuske, CR (2001) Phylogenetic Specificity and Reproducibility and New Method for Analysis of Terminal Restriction Fragment Profiles of 16S rRNA Genes from Bacterial Communities. Appl. Environ. Microbiol. 67: 190-197&amp;lt;/ref&amp;gt;. Pseudo TRFs, on the other hand, are reproducible peaks and are linear to the amount of DNA loaded. These peaks are thought to be the result of ssDNA annealing on to itself and creating double stranded random restriction sites which are later recognized by the restriction enzyme resulting in a terminal fragment which does not represent any genuine genetic variant. It has been suggested that applying a DNA [[exonuclease]] such as the Mung bean exonuclease prior to the digestion stage might eliminate such artifact.&lt;br /&gt;
&lt;br /&gt;
==Interpretation of data==&lt;br /&gt;
The data resulting from the electropherogram is normally interpreted in one of the following ways.&lt;br /&gt;
&lt;br /&gt;
===Pattern comparison=== &lt;br /&gt;
In pattern comparison the general shapes of electropherograms of different samples are compared for changes such as presence absence of peaks between treatments, their relative size etc.&lt;br /&gt;
&lt;br /&gt;
===Complementing with a clone library===&lt;br /&gt;
If a clone library is constructed in parallel to the T-RFLP analysis then the clones can be used to assess and interpret the T-RFLP profile. In this method the TRF of each clone is determined either directly (i.e. performing T-RFLP analysis on each single clone) or by ‘’in-silico’’ analysis of that clone’s sequence. By comparing the T-RFLP profile to a clone library it is possible to validate each of the peaks as genuine as well as to assess the relative abundance of each variant in the library.&lt;br /&gt;
&lt;br /&gt;
===Peak resolving using a database=== &lt;br /&gt;
Several computer applications attempt to relate the peaks in an electropherogram to specific bacteria in a database. Normally this type of analysis is done by simultaneously resolving several profiles of a single sample obtained with different restriction enzymes. The software then resolves the profile by attempting to maximize the matches between the peaks in the profiles and the entries in the database so that the number of peaks left without a matching sequence is minimal. The software withdraws from the database only those sequences which have their TRFs in all analyzed profiles.&lt;br /&gt;
&lt;br /&gt;
===Multivariate analysis===&lt;br /&gt;
A recently growing way to analyze T-RFLP profiles is use multivariate statistical methods to interpret the T-RFLP data &amp;lt;ref name=multi&amp;gt; Zaid Abdo et al., “Statistical Methods for Characterizing Diversity of Microbial Communities by Analysis of Terminal Restriction Fragment Length Polymorphisms of 16S rRNA Genes,” Environmental Microbiology 8, no. 5 (May 2006): 929-938. &amp;lt;/ref&amp;gt;. Usually the methods applied are those commonly used in ecology and especially in the study of biodiversity. Among them ordinations and [[cluster analysis]] are the most widely used.&lt;br /&gt;
In order to perform multivariate statistical analysis on T-RFLP data, the data must first be converted to table known as a “sample by species table“ which depicts the different samples (T-RFLP profiles) versus the species (T-RFS) with the height or area of the peaks as values.&lt;br /&gt;
&lt;br /&gt;
==Advantages and disadvantages==&lt;br /&gt;
As T-RFLP is a fingerprinting technique its advantages and drawbacks are often discussed in comparison with other similar technique, mostly DGGE.&lt;br /&gt;
&lt;br /&gt;
===Advantages===&lt;br /&gt;
The major advantage of T-RFLP is the use of an automated sequencer which gives highly reproducible results for repeated samples. Although the genetic profiles are not completely reproducible and several minor peaks which appear are irreproducible the overall shape of the electropherogram and the ratios of the major peaks are considered reproducible. The use of an automated sequencer which outputs the results in a digital numerical format also enables an easy way to store the data and compare different samples and experiments. The numerical format of the data can and has been used for relative (though not absolute) quantification and statistical analysis. Although sequence data cannot be retrieved of the T-RFLP profile, ‘’in-silico’’ assignment of the peaks to existing sequences is possible to a certain extent.&lt;br /&gt;
&lt;br /&gt;
===Drawbacks===&lt;br /&gt;
The fact that only the terminal fragments are being read means that any two distinct sequences which share a terminal restriction site will result in one peak only on the electropherogram and will be indistinguishable. Indeed, when T-RFLP is applied on a complex microbial community the result is often a compression of the total diversity to normally 20-50 distinct peaks only representing each an unknown number of distinct sequences. Although this phenomenon makes the T-RFLP results easier to handle it, naturally, introduces biases and oversimplification of the real diversity. Attempts to minimize (but not overcome) this problem are often done by applying several restriction enzymes and/ or labeling both primers with a different fluorescent dye. The inability to retrieve sequences from T-RFLP often leads to the need to construct and analyze one or more clone libraries in parallel to the T-RFLP analysis which adds to the effort and complicates and analysis. The possible appearance of false (pseudo) T-RFs, as discussed above, is yet another drawback. To handle this researchers often only consider peaks which can be affiliated to sequences in a clone library.&lt;br /&gt;
&lt;br /&gt;
==References==&lt;br /&gt;
&lt;br /&gt;
{{reflist}}&lt;br /&gt;
&lt;br /&gt;
==External links and references== &lt;br /&gt;
1.        [[http://rdp8.cme.msu.edu/html/t-rflp_jul02.html Improved Protocol for T-RFLP by Capillary Electrophoresis]]&lt;br /&gt;
&lt;br /&gt;
2.        [[http://www.ibest.uidaho.edu/tools/trflp_stats/index.php Statistical methods for characterizing diversity of microbial communities by analysis of terminal restriction fragment length polymorphisms of 16S rRNA genes.]]&lt;br /&gt;
&lt;br /&gt;
3.        [[http://www.oardc.ohio-state.edu/trflpfragsort/whatisfragsort.php FragSort]]: A software for ‘’in-silico’’ assignment of T-RFLP profiles from Ohio State University.&lt;br /&gt;
&lt;br /&gt;
4.        [[http://mica.ibest.uidaho.edu/trflp.php T-RFLP Analysis (APLAUS+)]]: Another ‚‘in-silico‘‘ assignment tool on the website of the Microbial Community Analysis project of Idaho state University&lt;br /&gt;
&lt;br /&gt;
[[Category:Molecular biology]]&lt;br /&gt;
[[Category:Biochemistry methods]]&lt;br /&gt;
[[Category:Laboratory techniques]]&lt;br /&gt;
[[Category:Electrophoresis]]&lt;/div&gt;</summary>
		<author><name>157.193.8.204</name></author>
	</entry>
	<entry>
		<id>https://ideawaza.com/index.php?title=Introduction_to_Statistics&amp;diff=17111</id>
		<title>Introduction to Statistics</title>
		<link rel="alternate" type="text/html" href="https://ideawaza.com/index.php?title=Introduction_to_Statistics&amp;diff=17111"/>
		<updated>2008-04-29T12:19:18Z</updated>

		<summary type="html">&lt;p&gt;157.193.92.242: /* Preamble */&lt;/p&gt;
&lt;hr /&gt;
&lt;div&gt;{{nav|School:Mathematics/Undergraduate/Probability and Statistics}}&lt;br /&gt;
&lt;br /&gt;
==Examples==&lt;br /&gt;
*46% of people polled enjoy vanilla, while 54% prefer chocolate (+/-4% margin of error).&lt;br /&gt;
*A school&#039;s graduation has increased by 2%.&lt;br /&gt;
*A couple has 4 boys, and they are pregnant again:  what is the chance of having another boy?&lt;br /&gt;
*88% of people questioned feel that it is humane to put stray animals to sleep.&lt;br /&gt;
&lt;br /&gt;
These are basic examples of statistics we see everyday, but do we really understand what they mean?  With the study of statistics, these &#039;facts&#039; that we hear everyday can hopefully become a little more clear.  &lt;br /&gt;
&lt;br /&gt;
== Preamble ==&lt;br /&gt;
Statistics is permeated by probability.  An understanding of basic probability is critical for the understanding of the basic mathematical underpinning of statistics. Strictly speaking the word &#039;statistics&#039; means one or more measures describing the characteristics of a population. We use the term here in a more idiomatic sense to mean everything to do with sampling and the establishment of population measures.&lt;br /&gt;
&lt;br /&gt;
Most statistical procedures use probability to make a statement about the relationship between the independent variables and the dependent variables.  Typically, the question one attempts to answer using statistics is that there is a relationship between two variables.  To demonstrate that there is a relationship the experimenter must show that when one variable changes the second variable changes and that the amount of change is more than would be likely from mere chance alone.&lt;br /&gt;
&lt;br /&gt;
There are two ways to figure the probability of an event.  The first is to do a mathematical calculation to determine how often the event can happen.  The second is to observe how often the event happens by counting the number of times the event could happen and also counting the number of times the event actually does happen.&lt;br /&gt;
&lt;br /&gt;
The use of a mathematical calculation is when a person can say that the chance of the event rolling a one on a six sided die is one in six.  The probability is figured by figuring the number of ways the event can happen and divide that number by the total number of possible outcomes.  Another example is in a well shuffled deck of cards, what is the probability of the event of drawing a three.  The answer is four in fifty two since there are four cards numbered three and there are a total of fifty two cards in a deck.  The chance of the event of drawing a card in the suite of diamonds is thirteen in fifty two (there are thirteen cards of each of the four suites).  The chance the event of drawing the three of diamonds is one in fifty two.&lt;br /&gt;
&lt;br /&gt;
Sometimes, the size of the total event space, the number of different possible events, is not known.  In that case, you will need to observe the event system and count the number of times the event actually happens versus the number of times it could happen but doesn&#039;t.&lt;br /&gt;
&lt;br /&gt;
For instance, a warranty for a coffee maker is a probability statement.  The manufacturer calculates that the probability the coffee maker will stop working before the warranty period ends is low.  The way such a warranty is calculated involves testing the coffee maker to calculate how long the typical coffee maker continues to function.  Then the manufacturer uses this calculation to specify a warranty period for the device.  The actual calculation of the coffee maker&#039;s life span is made by testing coffee makers and the parts that make up a coffee maker and then using probability to calculate the warranty period.&lt;br /&gt;
&lt;br /&gt;
== Experiments, Outcomes and Events ==&lt;br /&gt;
&lt;br /&gt;
The easiest way to think of probability is in terms of experiments and their potential outcomes.  Many examples can be drawn from everyday experience:  On the drive home from work, you can encounter a flat tire, or have an uneventful drive; the outcome of an election can include either a win by candidate A, B, or C, or a runoff.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Definition:&#039;&#039;&#039; The entire collection of possible outcomes from an experiment is termed the &#039;&#039;sample space&#039;&#039;, indicated as &#039;&#039;&#039;&amp;lt;math&amp;gt;\Omega&amp;lt;/math&amp;gt;&#039;&#039;&#039; (&#039;&#039;Omega&#039;&#039;)&lt;br /&gt;
&lt;br /&gt;
The simplest (albeit uninteresting) example would be an experiment with only one possible outcome, say &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;.  From elementary set theory, we can express the sample space as follows:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\Omega = \{ A \} &amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
A more interesting example is the result of rolling a six sided dice.  The sample space for this experiment is:&lt;br /&gt;
&lt;br /&gt;
&amp;lt;math&amp;gt;\Omega = \{ 1,2,3,4,5,6 \}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We may be interested in &#039;&#039;events&#039;&#039; in an experiment.&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Definition:&#039;&#039;&#039; An &#039;&#039;event&#039;&#039; is some subset of outcomes from the &#039;&#039;sample space&#039;&#039;&lt;br /&gt;
&lt;br /&gt;
In the dice example, events of interest might include&amp;lt;br&amp;gt;&lt;br /&gt;
a) the outcome is an even number&amp;lt;br&amp;gt;&lt;br /&gt;
b) the outcome is less than three&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
These events can be expressed in terms of the possible outcomes from the experiment: &amp;lt;br&amp;gt;&lt;br /&gt;
a) : &amp;lt;math&amp;gt; \{2,4,6\} &amp;lt;/math&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
b) : &amp;lt;math&amp;gt; \{ 1,2 \}&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
We can borrow definitions from set theory to express events in terms of outcomes.  Here is a refresher of some terminology, and some new terms that will be important later: &amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt; \cup &amp;lt;/math&amp;gt; represents the Union of two events&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:Venn_A_union_B.png]]&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt; \cap &amp;lt;/math&amp;gt; represents the Intersection of two events&amp;lt;br&amp;gt;&lt;br /&gt;
[[Image:Venn_A_intersect_B.svg]]&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;\{\cdots\}^{c}&amp;lt;/math&amp;gt; represents the complement of an event.  For instance, &amp;quot;the outcome is an even number&amp;quot; is the complement of &amp;quot;the outcome is an odd number&amp;quot; in the dice example.&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt; A \backslash B &amp;lt;/math&amp;gt; represents &#039;&#039;difference&#039;&#039;, that is, &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; &#039;&#039;but not&#039;&#039; &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;.  For example, we may be interested in the event of drawing the queen of spades from a deck of cards.  This can be expressed as the event of drawing a queen, but not drawing a queen of hearts, diamonds or clubs.&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;\varnothing&amp;lt;/math&amp;gt; or &amp;lt;math&amp;gt;\{\}&amp;lt;/math&amp;gt; represent an &#039;&#039;impossible event&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;\Omega&amp;lt;/math&amp;gt; represents a &#039;&#039;certain event&#039;&#039; &amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; are called &#039;&#039;disjoint events&#039;&#039; if &amp;lt;math&amp;gt;A\cap B = \varnothing&amp;lt;/math&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Probability ==&lt;br /&gt;
Now that we know what events are, we should think a bit about a way to express the likelihood of an event occurring.  The classical definition of probability comes from the following.  If we can perform our experiment over and over in a way that is repeatable, we can count the number of times that the experiment gives rise to event &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;.  We also keep track of the number of times that we perform the same experiment.  If we repeat the experiment a large enough number of times, we can express the probability of event &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; as follows:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;P(A) = \frac{N_{A}}{N} &amp;lt;/math&amp;gt; &amp;lt;br&amp;gt;&lt;br /&gt;
where &amp;lt;math&amp;gt;N_{A}&amp;lt;/math&amp;gt; is the number of times event &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; occurred, and &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt; is the number of times the experiment was repeated. Therefore the equation can be read as &amp;quot;the probability of event &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; equals the number of times event &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; occurs divided by the number of times the experiment was repeated (or the number of times event &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; &#039;&#039;could have&#039;&#039; occurred).&amp;quot;  As &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt; approaches infinity, the fraction above approaches the true probability of the event &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt;.  The value of &amp;lt;math&amp;gt;P(A)&amp;lt;/math&amp;gt; is clearly between 0 and 1.  If our event is the &#039;&#039;certain event&#039;&#039; &amp;lt;math&amp;gt;\Omega&amp;lt;/math&amp;gt;, then for each time we perform the experiment, the event &amp;lt;math&amp;gt;\Omega&amp;lt;/math&amp;gt; is observed; &amp;lt;math&amp;gt;N_{\Omega} = N&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;P(\Omega)=1&amp;lt;/math&amp;gt;.  If our event is the &#039;&#039;impossible event&#039;&#039; &amp;lt;math&amp;gt;\varnothing&amp;lt;/math&amp;gt;, we know &amp;lt;math&amp;gt;N_{\varnothing}=0&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt; P(\varnothing) = 0&amp;lt;/math&amp;gt;.&lt;br /&gt;
&lt;br /&gt;
If &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; are &#039;&#039;disjoint events&#039;&#039;, then whenever event &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; is observed, then it is impossible for event &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; to be observed simultaneously.  Therefore the number of times events &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; union &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; occurs are equal to the number of times event &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; occurred plus the number of times &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; occurs. This can be expressed as: &amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;N(A\cup B) = N(A) + N(B)&amp;lt;/math&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
Given our definition of probability, we can arrive at the following:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;P(A\cup B) = P(A) + P(B)&amp;lt;/math&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
At this point it&#039;s worth remembering that not all events are disjoint events.  For events that are not disjoint, we end up with the following probability definition.&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt; P(A\cup B) = P(A) + P(B) - P(A\cap B)&amp;lt;/math&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
How can we see this from example?  Well, let&#039;s consider drawing from a deck of cards.  I&#039;ll define two &#039;&#039;events&#039;&#039;: &amp;quot;drawing a Queen&amp;quot;, and &amp;quot;drawing a Spade&amp;quot;.  It is immediately clear that these are not disjoint events, because you can draw a queen that is also a spade.  There are four queens in the deck, so if we perform the experiment of drawing a card, putting it back in the deck and shuffling (what statisticians refer to as &#039;&#039;sampling with replacement&#039;&#039;, we will end up with a probability of &amp;lt;math&amp;gt;\frac{1}{13}&amp;lt;/math&amp;gt; for a queen draw.  By the same argument, we obtain a probability for drawing a spade as &amp;lt;math&amp;gt;\frac{1}{4}&amp;lt;/math&amp;gt;.  The expression &amp;lt;math&amp;gt;P(A\cup B)&amp;lt;/math&amp;gt; here can be translated as &amp;quot;the chance of drawing a queen or a spade&amp;quot;.  If we incorrectly assume that for this case &amp;lt;math&amp;gt;P(A\cup B) = P(A) + P(B)&amp;lt;/math&amp;gt;, we can simply add our probabilities together for &amp;quot;the chance of drawing a queen or a spade&amp;quot; as &amp;lt;math&amp;gt;\frac{1}{13}+\frac{1}{4}&amp;lt;/math&amp;gt;.  If we were to gather some data experimentally, we would find that our results would differ from the prediction -- the probability observed would be slightly less than &amp;lt;math&amp;gt;\frac{1}{13}+\frac{1}{4}&amp;lt;/math&amp;gt;.  Why?  Because we&#039;re counting the queen of spades twice in our expression, once as a spade, and again as a queen.  We need to count it only once, as it can only be drawn with probability of &amp;lt;math&amp;gt;\frac{1}{52}&amp;lt;/math&amp;gt;.  [[Introduction to Statistics/Still confused?|Still confused?]] &amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
&#039;&#039;&#039;Proof:&#039;&#039;&#039; If &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; are not disjoint, we have to avoid the double counting problem by exactly specifying their union.&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;A \cup B = A \cup (B \backslash A) &amp;lt;/math&amp;gt; so &amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;P(A \cup B) = P(A \cup (B \backslash A)) &amp;lt;/math&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B \backslash A&amp;lt;/math&amp;gt; are disjoint sets.  We can then use the definition of disjoint events from above to express our desired result:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;P(A \cup B) = P(A) + P(B \backslash A)&amp;lt;/math&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
We also know that&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;P(B \backslash A) = P(B) - P(B\cap A)&amp;lt;/math&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
so&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt; P(A \cup B) = P(A) + P(B) - P(B\cap A)&amp;lt;/math&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
Whew!  Our first proof.  I hope that wasn&#039;t too [[dry]].&lt;br /&gt;
&lt;br /&gt;
== Conditional Probability ==&lt;br /&gt;
Many events are conditional on the occurance of other events.  Sometimes this coupling is weak.  One event may become more or less probable depending on our knowledge that another event has occured.  For instance, the probability that your friends and relatives will call asking for money is likely to be higher if you win the lottery.  In my case, I don&#039;t think this probability would change.&lt;br /&gt;
&lt;br /&gt;
Let&#039;s get formal for a second and remember our original definition of probability.&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;P(A) = \frac{N_{A}}{N}&amp;lt;/math&amp;gt;&lt;br /&gt;
Consider an additional event &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;, and a situation where we are only interested in the probability of the occurance of &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; when &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; occurs.  A way at this probability is to perform a set of experiments (&#039;&#039;trials&#039;&#039;) and only record our results when the event &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; occurs.  In other words&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt; \frac{N_{A\cap B}}{N_{B}} &amp;lt;/math&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
We can divide through on top and bottom by &amp;lt;math&amp;gt;N&amp;lt;/math&amp;gt; the total number of trials to get &amp;lt;math&amp;gt;P(A\cap B)/P(B)&amp;lt;/math&amp;gt;.  We define this as &#039;&#039;&#039;&#039;conditional probability&#039;&#039;&#039;&#039;:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;P(A|B) = \frac{P(A\cap B)}{P(B)}&amp;lt;/math&amp;gt; &amp;lt;br&amp;gt;&lt;br /&gt;
which when spoken, takes the sound &amp;quot;probability of &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; given &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt;.&amp;quot;&lt;br /&gt;
&lt;br /&gt;
[[/A Totally Confusing Problem that Shows how Difficult it is to Conquer Intuition/]]&lt;br /&gt;
&lt;br /&gt;
=== Bayes&#039; Law ===&lt;br /&gt;
&#039;&#039;&#039;THIS THEOREM IS NOT GOOD, THERE IS AN ERROR !!!!!!!&#039;&#039;&#039;&amp;lt;br&amp;gt;&lt;br /&gt;
&#039;&#039;&#039;JUST SEE THE RIGHT ONE AT THIS PAGE:&#039;&#039;&#039; &amp;lt;br&amp;gt;&lt;br /&gt;
http://en.wikipedia.org/wiki/Bayes&#039;_theorem&amp;lt;br&amp;gt;&lt;br /&gt;
An important theorem in statistics is &#039;&#039;&#039;Bayes&#039; Law&#039;&#039;&#039;, which states that &amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt; P(B|A) = \frac{P(A|B)P(A)}{P(B)}&amp;lt;/math&amp;gt;,&amp;lt;br&amp;gt;&lt;br /&gt;
It is easy to prove.  We start with identical expressions for &amp;lt;math&amp;gt;P(A\cap B)&amp;lt;/math&amp;gt;.&amp;lt;br&amp;gt;&lt;br /&gt;
We know that: &amp;lt;math&amp;gt; P(A\cap B) = P(B\cap A) &amp;lt;/math&amp;gt;,&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt; P(A\cap B) = \frac{P(A|B)}{P(B)}&amp;lt;/math&amp;gt;, and&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt; P(B\cap A) = \frac{P(B|A)}{P(A)}&amp;lt;/math&amp;gt;.&amp;lt;br&amp;gt;&lt;br /&gt;
Since &amp;lt;math&amp;gt; P(A\cap B) = P(B\cap A) &amp;lt;/math&amp;gt;,&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;\frac{P(A|B)}{P(B)} = \frac{P(B|A)}{P(A)}&amp;lt;/math&amp;gt;.&amp;lt;br&amp;gt;&lt;br /&gt;
A Simple rearrangement of above line gives us &#039;&#039;&#039;Bayes&#039; Law&#039;&#039;&#039;.&lt;br /&gt;
&lt;br /&gt;
== Independence ==&lt;br /&gt;
&lt;br /&gt;
Two events &amp;lt;math&amp;gt;A&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;B&amp;lt;/math&amp;gt; are called &#039;&#039;independent&#039;&#039; if the occurence of one has absolutely no effect on the probability of the occurence of the other. Mathematically, this is expressed as:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;P(A\cap B) = P(A)P(B)&amp;lt;/math&amp;gt;.&amp;lt;br&amp;gt;&lt;br /&gt;
&lt;br /&gt;
== Random Variables ==&lt;br /&gt;
&lt;br /&gt;
It&#039;s usually possible to represent the outcome of experiments in terms of integers or real numbers.  For instance, in the case of conducting a poll, it becomes a little cumbersome to present the outcomes of each individual respondant.  Let&#039;s say we poll ten people for their voting preferences (Republican - R, or Democrat - D) in two different electorial districts.  Our results might look like this:&amp;lt;br&amp;gt;&lt;br /&gt;
&amp;lt;math&amp;gt;\{RRRDRRDRRR\}&amp;lt;/math&amp;gt; and &amp;lt;math&amp;gt;\{DDDDDRDDDD\}&amp;lt;/math&amp;gt;&amp;lt;br&amp;gt;&lt;br /&gt;
But we&#039;re probably only interested in the overall breakdown in voting preference for each district.  If we assign an integer value to each outcome, say 0 for Democrat and 1 for Republican, we can obtain a concise summary of voting preference by district simply by adding the results together.&lt;br /&gt;
&lt;br /&gt;
== Discrete and Continuous Random Variables ==&lt;br /&gt;
There are two important subclasses of random variables: discrete random variable (DRV) and continuous random variable (CRV).&lt;br /&gt;
Discrete random variables take only countably many values. It means that we can list the set of all possible values that a discrete random variable can take, or in other words, the number of possible values in the set that the variable can take is finite. If the possible values that a DRV X can take are a0,a1,a2,...an, the probability that X  takes each is p0=P(X=a0), p1=P(X=a1), p2=P(X=a2),...pn=P(X=an). All these probabilites are greater than or equal zero.&lt;br /&gt;
&lt;br /&gt;
For continuous random variables, we cannot list all possible values that a continuous variable can take because the number of values it can take is extremely large. It means that there is no use to calculate the probability of each value separately because the probability that the variable takes a particular value is extremely small and can be considered zero P(X=x)=0).&lt;br /&gt;
&lt;br /&gt;
== Distribution Functions ==&lt;br /&gt;
&lt;br /&gt;
== Expectation Values ==&lt;br /&gt;
&lt;br /&gt;
==See also==&lt;br /&gt;
* [[Topic:Statistics]]&lt;br /&gt;
* [[Introduction to research]]&lt;br /&gt;
* [[Statistical Economics]]&lt;br /&gt;
* [[Topic:Actuarial mathematics]]&lt;br /&gt;
* [[Introduction to Likelihood Theory]]&lt;br /&gt;
* [[Introduction to Classical Statistics]]&lt;br /&gt;
* [[Wikiversity:Statistics 202]]&lt;br /&gt;
* [[Introduction to probability and statistics]]&lt;br /&gt;
* [[Bayesian Statistics]]&lt;br /&gt;
* [[Statistics for Business Decisions]]&lt;br /&gt;
* [[Wikiversity:Statistics]]&lt;br /&gt;
&lt;br /&gt;
[[Category:Mathematics]]&lt;br /&gt;
[[Category:Statistics]]&lt;br /&gt;
[[Category:Introductions]]&lt;br /&gt;
&lt;br /&gt;
[[fr:Initiation à la statistique]]&lt;/div&gt;</summary>
		<author><name>157.193.92.242</name></author>
	</entry>
</feed>