Security scan is checking this page.

Chapter 1

Sampling and Data

Read chapter 1 in the book

Summary

The science of statistics deals with the collection, analysis, interpretation, and presentation of data . We see and use data in our everyday lives Have class members write down the average time—in hours, to the nearest half-hour—they sleep per night.

Key terms

data
a set of observations (a set of possible outcomes); most data can be put into two groups: qualitative (an attribute whose value is indicated by a label) or quantitative (an…
statistic
a numerical characteristic of the sample; a statistic estimates the corresponding population parameter
average
also called mean; a number that describes the central tendency of the data
cluster sampling
a method for selecting a random sample and dividing the population into groups (clusters); use simple random sampling to select a set of clusters; every individual in the chosen…
control group
a group in a randomized experiment that receives an inactive treatment but is otherwise managed exactly as the other groups
convenience sampling
a nonrandom method of selecting a sample; this method selects individuals that are easily accessible and may result in biased data
cumulative relative frequency
the term applies to an ordered set of observations from smallest to largest. The cumulative relative frequency is the sum of the relative frequencies for all values that are less…
sampling with replacement
once a member of the population is selected for inclusion in a sample, that member is returned to the population for the selection of the next individual

Chapter 2

Descriptive Statistics

Read chapter 2 in the book

Summary

Each data point in one data set is matched with exactly one point from the other set

Key terms

mean
a number that measures the central tendency of the data; a common name for mean is average . The term mean is a shortened form of arithmetic mean . By definition, the mean for a…
median
a number that separates ordered data into halves; half the values are the same number or smaller than the median, and half the values are the same number or larger than the…
paired data set
two data sets that have a one-to-one relationship so that both data sets are the same size, and each data point in one data set is matched with exactly one point from the other set
percentile
a number that divides ordered data into hundredths; percentiles may or may not be part of the data. The median of the data is the second quartile and the 50 th percentile The…
quartiles
the numbers that separate the data into quarters; quartiles may or may not be part of the data; the second quartile is the median of the data
skewed
used to describe data that is not symmetrical; when the right side of a graph looks chopped off compared to the left side, we say it is skewed to the left . When the left side of…
standard deviation
a number that is equal to the square root of the variance and measures how far data values are from their mean; notation: s for sample standard deviation and σ for population…

Chapter 3

Probability Topics

Read chapter 3 in the book

Summary

Probability is a measure that is associated with how certain we are of results, or outcomes, of a particular activity. If A and B are any two mutually exclusive events, then P ( A OR B ) = P ( A ) + P (B), and When the activity is a planned operation carried out under controlled conditions, it is called an experiment . If the result is not predetermined, then the experiment is said to be a chance experiment.

Key terms

probability
a number between zero and one, inclusive, that gives the likelihood that a specific event will occur; the foundation of statistics is given by the following three axioms (by A.N…
event
a subset of the set of all outcomes of an experiment; the set of all outcomes of an experiment is called a sample space and is usually denoted by S . An event is an arbitrary…
mutually exclusive
two events are mutually exclusive if the probability that they both happen at the same time is zero; if events A and B are mutually exclusive, then P ( A AND B ) = 0
activity
a planned operation carried out under controlled conditions, it is called an experiment
and B
any two mutually exclusive events, then P ( A OR B ) = P ( A ) + P (B), and
experiment
a planned activity carried out under controlled conditions

Chapter 4

Discrete Random Variables

Read chapter 4 in the book

Summary

The probability p of a success is the same for any trial (so the probability q = 1 - p of a failure is the same for any trial) The n trials are independent and are repeated using identical conditions There are one or more Bernoulli trials with all failures except the last one, which is a success There are only two possible outcomes called success and failure for each trial

Key terms

Bernoulli trials
an experiment with the following characteristics: There are only two possible outcomes called success and failure for each trial The probability p of a success is the same for…
probability p of a success
the same for any trial (so the probability q = 1 - p of a failure is the same for any trial)
n trials
independent and are repeated using identical conditions
binomial experiment
a statistical experiment that satisfies the following three conditions: There are a fixed number of trials, n There are only two possible outcomes, called success and, failure …
binomial probability distribution
a discrete random variable (RV) that arises from Bernoulli trials; there are a fixed number, n , of independent trials Independent means that the result of any trial (for…
expected value
expected arithmetic average when an experiment is repeated many times; also called the mean; notations μ ; for a discrete random variable (RV) with probability distribution…
geometric distribution
a discrete random variable (RV) that arises from the Bernoulli trials; the trials are repeated until the first success. The geometric variable X is defined as the number of…
geometric experiment
a statistical experiment with the following properties: There are one or more Bernoulli trials with all failures except the last one, which is a success In theory, the number of…

Chapter 5

Continuous Random Variables

Read chapter 5 in the book

Summary

Since the maximum probability is one, the maximum area is also one. We begin by defining a continuous probability density function. Intermediate algebra may have been your first formal introduction to functions.

Key terms

maximum probability
one, the maximum area is also one
decay parameter
The decay parameter describes the rate at which probabilities decay to zero for increasing values of x . It is the value m in the probability density function f ( x ) = me (- mx…
exponential distribution
a continuous random variable (RV) that appears when we are interested in the intervals of time between some random events, for example, the length of time between emergency…
Poisson distribution
a distribution function that gives the probability of a number of events occurring in a fixed interval of time or space if these events happen with a known average rate and…
uniform distribution
a continuous random variable (RV) that has equally likely outcomes over the domain, a < x < b . Notation— X ~ U ( a , b ). The mean is μ = a + b 2 and the standard deviation is σ…
conditional probability
the likelihood that an event will occur given that another event has already occurred

Chapter 6

The Normal Distribution

Read chapter 6 in the book

Summary

The standardized normal distribution is a type of normal distribution, with a mean of 0 and standard deviation of 1. Z -scores can be looked up in a Z -Table of Standard Normal Distribution, in order to find the area under the standard normal curve, between a score and the mean, between two scores, or above or below a score. The standard normal distribution allows us to interpret standardized scores and provides us with one table that we may use, in order to compute areas under the normal curve, for an infinite number of data sets, no… It represents a distribution of standardized scores, called z -scores , as opposed to raw scores (the actual data values).

Key terms

normal distribution
a continuous random variable (RV) where μ is the mean of the distribution and σ is the standard deviation; notation: X ~ N ( μ , σ ). If μ = 0 and σ = 1, the RV is called the…
standard normal distribution
a continuous random variable (RV) X ~ N (0, 1); when X follows the standard normal distribution, it is often noted as Z ~ N (0, 1)
standardized normal distribution
a type of normal distribution, with a mean of 0 and standard deviation of 1
z-score
the linear transformation of the form z = x - μ σ ; if this transformation is applied to any normal distribution X ~ N ( μ , σ ), the result is the standard normal distribution Z…

Chapter 7

The Central Limit Theorem

Read chapter 7 in the book

Summary

The sampling distribution of the mean approaches a normal distribution as n , the sample size , increases The central limit theorem for sample means says that if you keep drawing larger and larger samples (such as rolling one, two, five, and finally, ten dice) and calculating their means , the sample means form their own… The normal distribution has the same mean as the original distribution and a variance that equals the original variance divided by the sample size. The variable n is the number of values that are averaged together, not the number of times the experiment is done

Key terms

central limit theorem
given a random variable (RV) with a known mean, μ , and known standard deviation, σ , and sampling with size n , we are interested in two new RVs: the sample mean, X ¯ , and the…
average
a number that describes the central tendency of the data; there are a number of specialized averages, including the arithmetic mean, weighted mean, median, mode, and geometric mean
mean
a number that measures the central tendency; a common name for mean is average ; the term mean is a shortened form of arithmetic mean; . by definition, the mean for a sample…
normal distribution
a continuous random variable (RV) with probability density function (pdf) f ( x ) = 1 σ 2 π e - ( x - μ ) 2 σ 2 f ( x ) = 1 σ 2 π e - ( x - μ ) 2 σ 2 , where μ is the mean of the…
sampling distribution
given simple random samples of size n from a given population with a measured characteristic such as mean, proportion, or standard deviation for each sample, the probability…
variable n
the number of values that are averaged together, not the number of times the experiment is done
exponential distribution
a continuous random variable (RV) that appears when we are interested in the intervals of time between a random events; for example, the length of time between emergency arrivals…
uniform distribution
a continuous random variable (RV) that has equally likely outcomes over the domain a < x < b ; often referred as the rectangular distribution because the graph of the pdf has the…

Chapter 8

Confidence Intervals

Read chapter 8 in the book

Summary

Information that is known about the distribution (for example, known standard deviation), and The pdf is symmetrical about its mean of zero. There is a family of t -distributions: Each representative of the family is completely defined by the number of degrees of freedom, which is one less than the number of data It is continuous and assumes any real values

Key terms

standard deviation
a number that is equal to the square root of the variance and measures how far data values are from their mean; notation: s for sample standard deviation and σ for population…
family
completely defined by the number of degrees of freedom, which is one less than the number of data
pdf
symmetrical about its mean of zero
binomial distribution
a discrete random variable (RV) that arises from Bernoulli trials; there are a fixed number, n , of independent trials Independent means that the result of any trial (for…
confidence interval ( CI )
an interval estimate for an unknown population parameter. This depends on the following: the desired confidence level, information that is known about the distribution (for…
inferential statistics
also called statistical inference or inductive statistics; this facet of statistics deals with estimating a population parameter based on a sample statistic For example, if four…
Student's t -distribution
investigated and reported by William S. Gossett in 1908 and published under the pseudonym Student the major characteristics of the random variable (RV) are as follows: It is…

Chapter 9

Hypothesis Testing with One Sample

Read chapter 9 in the book

Summary

A bell-shaped continuous random variable X , with center at the mean value ( μ ) and distance from the center to the inflection points of the bell curve given by the standard deviation ( σ ) We write X ~ N ( μ , σ ) . If the mean value is 0 and the standard deviation is 1, the random variable is called the standard normal distribution, and it is denoted with the letter Z Information that is known about the distribution (for example, known standard deviation) The pdf is symmetrical about its mean of zero.

Key terms

hypothesis
a statement about the value of a population parameter; in the case of two hypotheses, the statement assumed to be true is called the null hypothesis (notation H 0 ) and the…
hypothesis testing
based on sample evidence, a procedure for determining whether the hypothesis stated is a reasonable statement and should not be rejected, or is unreasonable and should be rejected
standard deviation
a number that is equal to the square root of the variance and measures how far data values are from their mean; notation: s for sample standard deviation and σ for population…
mean value
0 and the standard deviation is 1, the random variable is called the standard normal distribution, and it is denoted with the letter Z
family
completely defined by the number of degrees of freedom, which is one less than the number of data items
pdf
symmetrical about its mean of zero
binomial distribution
a discrete random variable (RV) that arises from Bernoulli trials; there are a fixed number, n , of independent trials Independent means that the result of any trial (for…
confidence interval ( CI )
an interval estimate for an unknown population parameter This depends on the following: The desired confidence level. Information that is known about the distribution (for…

Chapter 10

Hypothesis Testing with Two Samples

Read chapter 10 in the book

Summary

The domain of the random variable (RV) is not necessarily a numerical set; the domain may be expressed in words; for example, if X = hair color, then the domain is {black, blond, gray, green, orange} We can tell what specific value x of the random variable X takes only after performing the experiment

Key terms

standard deviation
a number that is equal to the square root of the variance and measures how far data values are from their mean; notation: s for sample standard deviation and σ for population…
variable (random variable)
a characteristic of interest in a population being studied. Common notation for variables are uppercase Latin letters X , Y , Z ,... Common notation for a specific value from the…
pooled proportion
estimate of the common value of p 1 and p 2

Chapter 11

The Chi-Square Distribution

Read chapter 11 in the book

Summary

For the χ 2 distribution, the population mean is μ = df , and the population standard deviation is σ = 2 ( d f ) The random variable is shown as χ 2 , but it may be any uppercase letter The random variable for a chi-square distribution with k degrees of freedom is the sum of k independent, squared standard normal variables is The notation for the chi-square distribution is

Key terms

population mean
μ = df , and the population standard deviation is σ = 2 ( d f )
random variable
shown as χ 2 , but it may be any uppercase letter
contingency table
a table that displays sample values for two different factors that may be dependent or contingent on each other; facilitates determining conditional probabilities

Chapter 12

Linear Regression and Correlation

Read chapter 12 in the book

Summary

The variable x is the independent variable ; y is the dependent variable . The rate for services is $32 per hour plus a $31.50 one-time charge. Linear regression for two variables is based on a linear equation with one independent variable. Typically, you choose a value to substitute for the independent variable and then solve for the dependent variable

Key terms

variable x
the independent variable ; y is the dependent variable
rate for services
$32 per hour plus a $31.50 one-time charge
coefficient of correlation
a measure developed by Karl Pearson during the early 1900s that gives the strength of association between the independent variable and the dependent variable; r = n ∑ ​ x y - [ ∑…
outlier
an observation that does not fit the rest of the data

Chapter 13

F Distribution and One-way Anova

Read chapter 13 in the book

Summary

The test statistic for analysis of variance is the F ratio Samples (not necessarily of the same size) are randomly and independently selected from each population Samples (not necessarily of the same size) are randomly and independently selected from each population, and All populations of interest are normally distributed,

Key terms

one-way ANOVA
a method of testing whether the means of three or more populations are equal; the method is applicable if all populations of interest are normally distributed, the populations…
analysis of variance
also referred to as ANOVA; a method of testing whether the means of three or more populations are equal The method is applicable if all populations of interest are normally…
variance
mean of the squared deviations from the mean; the square of the standard deviation For a set of data, a deviation can be represented as x - x ¯ where x is a value of the data and…
same size)
randomly and independently selected from each population

Summaries and key terms on this page are taken from that chapter’s material already kept for this desk. They follow the OpenStax book. Margins is not affiliated with OpenStax. Resources, policy, and site safety