Security scan is checking this page.

Chapter 1

Sampling and Data

Read chapter 1 in the book

Summary

After you have studied probability and probability distributions, you will use formal methods for drawing conclusions from "good" data. The science of statistics deals with the collection, analysis, interpretation, and presentation of data . Organizing and summarizing data is called descriptive statistics . Two ways to summarize data are by graphing and by using numbers (for example, finding an average).

Key terms

Data
a set of observations (a set of possible outcomes); most data can be put into two groups: qualitative (an attribute whose value is indicated by a label) or quantitative (an…
Statistic
a numerical characteristic of the sample; a statistic estimates the corresponding population parameter
Probability
a number between zero and one, inclusive, that gives the likelihood that a specific event will occur
Average
also called mean or arithmetic mean; a number that describes the central tendency of the data
formal methods
called inferential statistics
Cluster Sampling
a method for selecting a random sample and dividing the population into groups (clusters); use simple random sampling to select a set of clusters. Every individual in the chosen…
Control Group
a group in a randomized experiment that receives an inactive treatment but is otherwise managed exactly as the other groups
Convenience Sampling
a nonrandom method of selecting a sample; this method selects individuals that are easily accessible and may result in biased data

Chapter 2

Descriptive Statistics

Read chapter 2 in the book

Summary

One simple graph, the stem-and-leaf graph or stemplot , comes from the field of exploratory data analysis. It is a good choice when the data sets are small. To create the plot, divide each observation of data into a stem and a leaf. The leaf consists of a final significant digit .

Key terms

Mean (arithmetic)
a number that measures the central tendency of the data; a common name for mean is 'average.' The term 'mean' is a shortened form of 'arithmetic mean.' By definition, the mean…
Median
a number that separates ordered data into halves; half the values are the same number or smaller than the median and half the values are the same number or larger than the…
Percentile
a number that divides ordered data into hundredths; percentiles may or may not be part of the data. The median of the data is the second quartile and the 50 th percentile. The…
Quartiles
the numbers that separate the data into quarters; quartiles may or may not be part of the data. The second quartile is the median of the data
Standard Deviation
a number that is equal to the square root of the variance and measures how far data values are from their mean; notation: s for sample standard deviation and σ for population…
Variance
mean of the squared deviations from the mean, or the square of the standard deviation; for a set of data, a deviation can be represented as x - where x is a value of the data and…
Frequency Table
a data representation in which grouped data is displayed along with the corresponding frequencies

Chapter 3

Probability Topics

Read chapter 3 in the book

Summary

If A and B are any two mutually exclusive events, then P ( A ∪ B ) = P ( A ) + P (B) Probability is a measure that is associated with how certain we are of outcomes of a particular experiment or activity. Three ways to represent a sample space are: to list the possible outcomes, to create a tree diagram, or to create a Venn diagram. If the result is not predetermined, then the experiment is said to be a chance experiment.

Key terms

Probability
a number between zero and one, inclusive, that gives the likelihood that a specific event will occur; the foundation of statistics is given by the following 3 axioms (by A.N…
Event
a subset of the set of all outcomes of an experiment; the set of all outcomes of an experiment is called a sample space and is usually denoted by S . An event is an arbitrary…
Mutually Exclusive
Two events are mutually exclusive if the probability that they both happen at the same time is zero. If events A and B are mutually exclusive, then P ( A ∩ B ) = 0
Tree Diagram
the useful visual representation of a sample space and events in the form of a “tree” with branches marked by possible outcomes together with associated probabilities…
Venn Diagram
the visual representation of a sample space and events in the form of circles or ovals showing their intersections
and B
any two mutually exclusive events, then P ( A ∪ B ) = P ( A ) + P (B)

Chapter 4

Discrete Random Variables

Read chapter 4 in the book

Summary

The probability p of a success is the same for any trial (so the probability q = 1 - p of a failure is the same for any trial) The n trials are independent and are repeated using identical conditions There are one or more Bernoulli trials with all failures except the last one, which is a success There are only two possible outcomes called “success” and “failure” for each trial

Key terms

Bernoulli Trials
an experiment with the following characteristics: There are only two possible outcomes called “success” and “failure” for each trial. The probability p of a success is the same…
probability p of a success
the same for any trial (so the probability q = 1 - p of a failure is the same for any trial)
n trials
independent and are repeated using identical conditions
Binomial Experiment
a statistical experiment that satisfies the following three conditions: There are a fixed number of trials, n . There are only two possible outcomes, called "success" and…
Binomial Probability Distribution
a discrete random variable (RV) that arises from Bernoulli trials; there are a fixed number, n , of independent trials. “Independent” means that the result of any trial (for…
Geometric Distribution
a discrete random variable (RV) that arises from the Bernoulli trials; the trials are repeated until the first success. The geometric variable X is defined as the number of…
Geometric Experiment
a statistical experiment with the following properties: There are one or more Bernoulli trials with all failures except the last one, which is a success. In theory, the number of…
Hypergeometric Experiment
a statistical experiment with the following properties: You take samples from two groups. You are concerned with a group of interest, called the first group. You sample without…

Chapter 5

Continuous Random Variables

Read chapter 5 in the book

Summary

Again with the Poisson distribution in Random Discrete Variables , the graph in Example 4.15 used boxes to represent the probability of specific values of the random variable. In this case, we were being a bit casual because the random variables of a Poisson distribution are discrete, whole numbers, and a box has width. The graph of a continuous probability distribution is a curve. Probability is represented by area under the curve.

Key terms

Poisson distribution
If there is a known average of μ events occurring per unit time, and these events are independent of each other, then the number of events X occurring in one unit of time has the…
decay parameter
The decay parameter describes the rate at which probabilities decay to zero for increasing values of x . It is the value m in the probability density function f ( x ) = me (- mx…
Exponential Distribution
a continuous random variable (RV) that appears when we are interested in the intervals of time between some random events, for example, the length of time between emergency…
Uniform Distribution
a continuous random variable (RV) that has equally likely outcomes over the domain, a < x < b ; it is often referred as the rectangular distribution because the graph of the pdf…
Conditional Probability
the likelihood that an event will occur given that another event has already occurred

Chapter 6

The Normal Distribution

Read chapter 6 in the book

Summary

The standard normal distribution is a normal distribution of standardized values called z -scores . The mean for the standard normal distribution is zero, and the standard deviation is one. The value x in the given equation comes from a known normal distribution with known mean μ and known standard deviation σ . The z -score tells how many standard deviations a particular x is away from the mean

Key terms

Normal Distribution
a continuous random variable (RV) with pdf f ( x ) = 1 σ 2 π e - ( x - μ ) 2 σ 2 , where μ is the mean of the distribution and σ is the standard deviation; notation: X ~ N ( μ …
Standard Normal Distribution
a continuous random variable (RV) X ~ N (0, 1); when X follows the standard normal distribution, it is often noted as Z ~ N (0, 1)
z-score
the linear transformation of the form z = x - μ σ or written as z = | x - μ | σ ; if this transformation is applied to any normal distribution X ~ N ( μ , σ ) the result is the…

Chapter 7

The Central Limit Theorem

Read chapter 7 in the book

Summary

Each sample mean is then treated like a single observation of this new distribution, the sampling distribution. The sampling distribution is a theoretical distribution. The genius of thinking this way is that it recognizes that when we sample we are creating an observation and that observation must come from some particular distribution. If this is discovered, then we can treat a sample mean just like any other observation and calculate probabilities about what values it might take on.

Key terms

Central Limit Theorem
Given a random variable with known mean μ and known standard deviation, σ , we are sampling with size n , and we are interested in two new RVs: the sample mean, X - . If the size…
Mean
a number that measures the central tendency; a common name for mean is "average." The term "mean" is a shortened form of "arithmetic mean." By definition, the mean for a sample…
Sampling Distribution
Given simple random samples of size n from a given population with a measured characteristic such as mean, proportion, or standard deviation for each sample, the probability…
genius of thinking this way
that it recognizes that when we sample we are creating an observation and that observation must come from some particular distribution
Average
a number that describes the central tendency of the data; there are a number of specialized averages, including the arithmetic mean, weighted mean, median, mode, and geometric mean
Finite Population Correction Factor
adjusts the variance of the sampling distribution if the population is known and more than 5% of the population is being sampled
Normal Distribution
a continuous random variable with pdf f ( x ) = 1 σ 2 π e - ( x - μ ) 2 σ 2 f ( x ) = 1 σ 2 π e - ( x - μ ) 2 σ 2 , where μ is the mean of the distribution and σ is the standard…
Standard Error of the Mean
the standard deviation of the distribution of the sample means, or σ n

Chapter 8

Confidence Intervals

Read chapter 8 in the book

Summary

Information that is known about the distribution (for example, known standard deviation), The pdf is symmetrical about its mean of zero It approaches the standard normal distribution as n get larger There is a "family" of t-distributions: each representative of the family is completely defined by the number of degrees of freedom, which depends upon the application for which the t is being used

Key terms

Normal Distribution
a continuous random variable (RV) with pdf f ( x ) = 1 σ 2 π e - ( x - μ ) 2 / 2 σ 2 f ( x ) = 1 σ 2 π e - ( x - μ ) 2 / 2 σ 2 , where μ is the mean of the distribution and σ is…
Standard Deviation
a number that is equal to the square root of the variance and measures how far data values are from their mean; notation: s for sample standard deviation and σ for population…
family
completely defined by the number of degrees of freedom, which depends upon the application for which the t is being used
pdf
symmetrical about its mean of zero
Binomial Distribution
a discrete random variable (RV) which arises from Bernoulli trials; there are a fixed number, n , of independent trials. “Independent” means that the result of any trial (for…
Confidence Interval (CI)
an interval estimate for an unknown population parameter. This depends on: the desired confidence level, information that is known about the distribution (for example, known…
Inferential Statistics
also called statistical inference or inductive statistics; this facet of statistics deals with estimating a population parameter based on a sample statistic. For example, if four…
Student's t -Distribution
investigated and reported by William S. Gossett in 1908 and published under the pseudonym Student; the major characteristics of this random variable (RV) are: It is continuous…

Chapter 9

Hypothesis Testing with One Sample

Read chapter 9 in the book

Summary

Information that is known about the distribution (for example, known standard deviation) The pdf is symmetrical about its mean of zero. However, it is more spread out and flatter at the apex than the normal distribution It approaches the standard normal distribution as n gets larger

Key terms

Hypothesis
a statement about the value of a population parameter, in case of two hypotheses, the statement assumed to be true is called the null hypothesis (notation H 0 ) and the…
Hypothesis Testing
Based on sample evidence, a procedure for determining whether the hypothesis stated is a reasonable statement and should not be rejected, or is unreasonable and should be rejected
Normal Distribution
a continuous random variable (RV) with pdf f ( x ) = 1 σ 2 π e - ( x - μ ) 2 σ 2 f ( x ) = 1 σ 2 π e - ( x - μ ) 2 σ 2 , where μ is the mean of the distribution, and σ is the…
Standard Deviation
a number that is equal to the square root of the variance and measures how far data values are from their mean; notation: s for sample standard deviation and σ for population…
family
completely defined by the number of degrees of freedom which is one less than the number of data items
pdf
symmetrical about its mean of zero
Binomial Distribution
a discrete random variable (RV) that arises from Bernoulli trials. There are a fixed number, n , of independent trials. “Independent” means that the result of any trial (for…
Central Limit Theorem
Given a random variable (RV) with known mean μ and known standard deviation σ. We are sampling with size n and we are interested in two new RVs - the sample mean, X ¯ . If the…

Chapter 10

Hypothesis Testing with Two Samples

Read chapter 10 in the book

Summary

The comparison of two independent population means is very common and provides a way to test the hypothesis that the two groups differ from each other. An observed difference between two sample means depends on both the means and the sample standard deviations. Very different means can occur by chance if there is great variation among the individual samples. The test statistic will have to account for this fact.

Key terms

Cohen’s d
a measure of effect size based on the differences between two means. If d is between 0 and 0.2 then the effect is small. If d approaches is 0.5, then the effect is medium, and if…
Independent Groups
two samples that are selected from two populations, and the values from one population are not related in any way to the values from the other population
Matched Pairs
two samples that are dependent. Differences between a before and after scenario are tested by testing one population mean of differences
Pooled Variance
a weighted average of two variances that can then be used when calculating standard error

Chapter 11

The Chi-Square Distribution

Read chapter 11 in the book

Summary

For the χ 2 distribution, the population mean is μ = df and the population standard deviation is σ = 2 ( d f ) The curve is nonsymmetrical and skewed to the right There is a different chi-square curve for each df . The notation for the chi-square distribution is

Key terms

population mean
μ = df and the population standard deviation is σ = 2 ( d f )
curve
nonsymmetrical and skewed to the right
Contingency Table
a table that displays sample values for two different factors that may be dependent or contingent on one another; it facilitates determining conditional probabilities
Goodness-of-Fit
a hypothesis test that compares expected and observed values in order to look for significant differences within one non-parametric variable. The degrees of freedom used equals…
Test for Homogeneity
a test used to draw a conclusion about whether two populations have the same distribution. The degrees of freedom used equals the (number of columns - 1)
Test of Independence
a hypothesis test that compares expected and observed values for contingency tables in order to test for independence between two variables. The degrees of freedom used equals…

Chapter 12

F Distribution and One-Way ANOVA

Read chapter 12 in the book

Summary

The test statistic for analysis of variance is the F -ratio Samples (not necessarily of the same size) are randomly and independently selected from each population All populations of interest are normally distributed The populations have equal standard deviations

Key terms

One-Way ANOVA
a method of testing whether or not the means of three or more populations are equal; the method is applicable if: all populations of interest are normally distributed. the…
Analysis of Variance
also referred to as ANOVA, is a method of testing whether or not the means of three or more populations are equal. The method is applicable if: all populations of interest are…
Variance
mean of the squared deviations from the mean; the square of the standard deviation. For a set of data, a deviation can be represented as x - where x is a value of the data and x…
same size)
randomly and independently selected from each population

Chapter 13

Linear Regression and Correlation

Read chapter 13 in the book

Summary

Perhaps unnoticed, all the data we have been using is for a single variable. The type of data described in the examples above and for any model of cause and effect is bivariate data — "bi" for two variables. In reality, statisticians use multivariate data, meaning many variables As we begin this section we note that the type of data we will be working with has changed.

Key terms

Linear
a model that takes data and regresses it into a straight line equation
Bivariate
two variables are present in the model where one is the “cause” or independent variable and the other is the “effect” of dependent variable
Multivariate
a system or model where more than one independent variable is being used to predict an outcome. There can only ever be one dependent variable, but there is no limit to the number…
Sum of Squared Errors (SSE)
the calculated value from adding up all the squared residual terms. The hope is that this value is very small when creating a model

Summaries and key terms on this page are taken from that chapter’s material already kept for this desk. They follow the OpenStax book. Margins is not affiliated with OpenStax. Resources, policy, and site safety