Unit 2 - Probability, Random Variables, and Probability Distributions
Built for AP Stats
Calculator drills
Practice with exam-style Desmos (College Board testing version) or a TI-84 for handheld paths — 9 decks, 36 drills.
Video Lessons
Table of Contents
2.1 - Tabular and Graphical Representations for the Distributions of Two Categorical Variables
Study Tips
- Visual vs. Numerical: Tables are the source of truth for raw counts, but graphs reveal the story of the relationship at a glance.
- Choose the Right Graph: Use side-by-side bar charts for raw frequency comparison and segmented or mosaic plots to compare proportions when group sizes differ.
Key Terms & Definitions
Bivariate Categorical Data
Data involving two distinct categorical variables measured on the same observational units.
Two-Way Table (Contingency Table)
A table used to organize frequencies or relative frequencies for two categorical variables, with one variable forming rows and the other forming columns.
Cell
The intersection of a row and a column in a two-way table, containing the count of individuals fitting both categories.
Side-by-Side Bar Chart
A graph grouping bars for each category to allow direct comparison of frequencies across levels of another variable.
Segmented Bar Chart
A chart where bars are stacked into a single bar representing 100% of the category, allowing for comparison of proportions regardless of group size.
Mosaic Plot
A specialized segmented bar chart where bar width is proportional to the sample size of the group.
Association
A relationship identified when the distribution of one categorical variable differs across the levels of the other.
2.2 - Summary Statistics for Two Categorical Variables
Study Tips
- The "Given" Denominator: When calculating conditional relative frequencies, the total of the given group (the row or column) is your denominator.
- Justification: Don't just claim an association exists; cite the specific percentage difference between groups to back it up.
Key Terms & Definitions
Joint Relative Frequency
The ratio of a specific cell frequency to the total count for the entire table.
Marginal Relative Frequency
A row or column total divided by the total for the entire table.
Conditional Relative Frequency
A relative frequency computed by restricting focus to a specific category; a cell frequency divided by its specific row or column total.
Independence
A state where the distribution of one categorical variable is identical across all levels of the other; the variables are not associated.
Justification
Using specific summary statistics as evidence to defend a claim about the variables in context.
2.3 - Estimating Probabilities Using Simulation
Study Tips
- The "Gold Standard": Simulations require more trials to reach accuracy. Always note that more trials reduce the impact of random luck.
- Structure: Follow the steps: assign digits, describe the trial, record counts, and calculate the relative frequency.
Key Terms & Definitions
Random Process
Any procedure that generates results determined by chance.
Outcome
The specific result of one trial of a random process.
Event
A collection of specific outcomes.
Simulation
A model of a random process used to estimate probabilities that are difficult to calculate directly.
Long-Run Relative Frequency
The probability of an event, defined as its relative frequency over a very large number of trials.
Law of Large Numbers (LLN)
The principle stating that as the number of independent trials increases, the empirical relative frequency converges toward the true probability.
Empirical Data
Data determined from actual observations or simulations.
2.4 - Introduction to Probability
Study Tips
- Complement Shortcut: If you need to find the probability of "at least one," calculate 1 − P(none).
- Sanity Check: Theoretical probabilities must always sum to 1 and be between 0 and 1 inclusive.
Key Terms & Definitions
Sample Space (S)
The set of all possible, non-overlapping outcomes.
Theoretical Probability (P(E))
The ratio of the number of outcomes in event E to the total number of outcomes in the sample space (when outcomes are equally likely).
Complement (E^C or E')
The event consisting of all outcomes in the sample space not in E.
Complement Rule
The calculation P(E^C) = 1 − P(E).
2.5 - Mutually Exclusive Events
Study Tips
- The Litmus Test: To prove two events are mutually exclusive, you must show the joint probability P(A ∩ B) = 0.
- Venn Visuals: If circles overlap, they are not mutually exclusive.
Key Terms & Definitions
Mutually Exclusive (Disjoint)
Events that cannot occur at the same time.
Joint Probability
The probability of the intersection of two events (P(A ∩ B)).
Intersection
The outcomes common to both events.
2.6 - Conditional Probability
Study Tips
- Denominator Watch: The event after the "given" bar (|) is always the denominator.
- Sequence Logic: Think of the General Multiplication Rule as a two-step story: P(A) · P(B|A) is the probability of A happening, then B happening given that A has already occurred.
Key Terms & Definitions
Conditional Probability (P(A|B))
The probability that event A will occur, given that B has already occurred.
Conditional Formula
P(A|B) = P(A ∩ B) / P(B).
General Multiplication Rule
The formula for the probability that both A and B occur: P(A ∩ B) = P(A) · P(B|A).
Dependent Events
Events where the outcome of one event changes the probability of the other.
2.7 - Independent Events and Unions of Events
Key Terms & Definitions
Independent Events
Two events are independent if the occurrence of one does not change the probability of the other.
Multiplication Rule for Independence
If events A and B are independent, the joint probability is the product of their individual probabilities: P(A ∩ B) = P(A) · P(B).
Union of Events (A ∪ B)
The event that A occurs, B occurs, or both occur.
General Addition Rule
The formula for finding the probability of a union: P(A ∪ B) = P(A) + P(B) − P(A ∩ B).
2.8 - Introduction to Random Variables and Probability Distributions
Key Terms & Definitions
Random Variable
A variable whose values are numerical outcomes resulting from a random phenomenon.
Discrete Random Variable
A random variable that can only take on a countable or finite number of values.
Probability Distribution
A representation (graph, table, or function) showing the probability associated with every possible value of a discrete random variable.
Sum of Probabilities Property
The fundamental rule that the sum of the probabilities over all possible values of a discrete random variable must equal 1.
Cumulative Probability Distribution
A representation showing the probability of the random variable being less than or equal to each value.
Random Phenomenon
A process that generates numerical outcomes determined by chance.
Construction Methods
Discrete probability distributions can be determined through the rules of probability or estimated via simulation.
2.9 - Parameters of Random Variables
Key Terms & Definitions
Parameter
A numerical value that measures a characteristic of a probability distribution of a random variable or a population.
Fixed Value
The property that a parameter is a single, constant value, unlike a statistic which varies from sample to sample.
Expected Value (Mean)
A parameter denoted by E(X) or μ_X that represents the long-run average outcome of a random variable.
- •Formula: μ_X = Σ x_i · P(x_i).
Expected Value Formula
Calculated as μ_X = Σ x_i · P(x_i), where x_i is a possible value and P(x_i) is its probability.
Standard Deviation of a Distribution
A parameter denoted by SD(X) or σ_X that measures the typical deviation of the values of the random variable from the mean.
- •Formula: σ_X = √(Σ(x_i − μ_X)² · P(x_i)).
Standard Deviation Formula
Calculated as σ_X = √(Σ(x_i − μ_X)² · P(x_i)).
Variance
The square of the standard deviation of a random variable, denoted as V(X) or σ_X².
2.10 - The Binomial Distribution
Key Terms & Definitions
Binomial Random Variable
A discrete random variable that counts the number of successes in a fixed number of independent trials.
Success/Failure
The two possible, mutually exclusive outcomes of a single trial in a binomial setting.
Independence of Trials
The requirement that the outcome of one trial does not affect the probability of success in any other trial.
Binomial Mean (μ_X)
The parameter for the expected number of successes, calculated as np.
Binomial Standard Deviation (σ_X)
The parameter measuring the spread of the number of successes, calculated as √(np(1 − p)).
Binomial Probability Function
The formula P(X = x) = (n choose x) p^x (1 − p)^(n−x), used to find the probability of exactly x successes in n trials.
Binomial Parameters
The specific values n (number of trials) and p (probability of success) that define a unique binomial distribution.
2.11 - The Normal Distribution
Key Terms & Definitions
Continuous Random Variable
A variable that takes on any value in a domain; probabilities are associated with intervals.
Normal Distribution
A continuous, unimodal, bell-shaped, and symmetric model.
Normal Curve Parameters
The mean (μ) and standard deviation (σ).
Standard Normal Distribution
A normal distribution where μ = 0 and σ = 1.
Empirical Rule (68-95-99.7 Rule)
A heuristic for estimating areas under the normal curve.
Area Under the Curve
The representation of probability; the total area under the curve is always 1.
Z-score
A standardized score indicating distance from the mean in units of standard deviation.
Percentile
A measure of relative position based on the proportion of values below a given point.
2.12 - Sampling Distributions and the Central Limit Theorem
Key Terms & Definitions
Sampling Distribution of a Statistic
The distribution of all possible values of a statistic for samples of a given size from a population.
Simulation of Sampling Distribution
Using random sampling to approximate a sampling distribution.
Randomization Distribution
A distribution generated by repeatedly reallocating response values to treatment groups to approximate a sampling distribution.
Central Limit Theorem (CLT)
The statement that the sampling distribution of a sample mean is approximately normal for sufficiently large samples.
Sample Size Effect
The property that a larger sample size improves the normal approximation provided by the CLT.