Unit 5 - Regression Analysis
Built for AP Stats
Calculator drills
Practice with exam-style Desmos (College Board testing version) or a TI-84 for handheld paths — 9 decks, 36 drills.
Table of Contents
5.1 - Graphical Representations Between Two Quantitative Variables
Key Terms & Definitions
Bivariate quantitative data
A data set consisting of observations of two different quantitative variables measured on the same individuals in a sample or population.
- •Formed by coordinate pairs (x, y) where both metrics capture numerical counts or physical scales.
Cluster
Concentrations of data usually separated by gaps in a distribution.
- •Clusters often indicate the presence of hidden categorical subgroups within the broader dataset.
Direction
The type of association between two variables in a scatter plot, described as positive or negative.
- •Classified strictly as positive or negative based on the upward or downward slant of the data cloud.
Explanatory variable
A variable whose values are used to explain or predict corresponding values for the response variable.
- •Its values are treated as the given conditions used to explain variance in the accompanying response metric.
Form
The pattern or shape of the relationship between two variables in a scatter plot, such as linear or non-linear.
- •Categorized broadly as linear or non-linear depending on whether the rate of change is constant or curved.
Linear
A form of association in a scatter plot where the points follow a straight-line pattern.
- •Avoid saying "the points form a line"; instead, explain that the trend advances at a fixed, uniform rate.
Negative association
A relationship between two variables where as values of one variable increase, values of the other variable tend to decrease.
- •On a graph, a negative trend shows an overarching data track that paths downward from left to right.
Non-linear
A form of association in a scatter plot where the points do not follow a straight-line pattern.
- •Standard simple linear equations are inappropriate for optimizing or tracking curved bivariate shapes.
Outlier
Data points that are unusually small or large relative to the rest of the data.
- •Can drastically manipulate distance averages and model tracks depending on its specific placement.
Positive association
A relationship between two variables where as values of one variable increase, values of the other variable tend to increase.
- •Visually recognized as a coordinate cloud that slants consistently upward from left to right.
Response variable
A variable whose values are being explained or predicted based on the explanatory variable.
- •The dependent variable whose internal variation is the main property we seek to explain.
Scatter plot
A graph that displays the relationship between two quantitative variables using points plotted on a coordinate plane.
- •The indispensable initial tool used to diagnose form, direction, strength, and anomalies.
Strength
A measure of how closely individual points in a scatter plot follow a specific pattern, described as strong, moderate, or weak.
- •Linguistically classified as strong, moderate, or weak depending on the amount of scattered noise.
5.2 - Correlation
Key Terms & Definitions
Causation
A relationship where changes in one variable directly cause changes in another variable.
- •Can never be verified by observing high correlation alone; requires well-designed, randomized experiments.
Correlation
A numerical measure (r) that describes the strength and direction of a linear relationship between two variables, ranging from -1 to 1.
- •Sign tracks the direction, while absolute proximity to 1 validates the relative strength of the linear model.
Linear model
A mathematical representation of the linear relationship between two variables.
- •Employed to standardize, interpret, and calculate estimated behaviors across a linear domain space.
Linear relationship
A relationship between two variables that can be described by a straight line.
- •The singular target pattern designed to be evaluated by the correlation metric r.
Quantitative variable
A variable that is measured numerically and can take on a range of values, allowing for mathematical operations and statistical analysis.
- •Must possess measurable units, separating it from group label definitions like zip codes or names.
5.3 - Linear Regression Models
Key Terms & Definitions
Explanatory variable
A variable whose values are used to explain or predict corresponding values for the response variable.
- •Its values are treated as the given conditions used to explain variance in the accompanying response metric.
Response variable
A variable whose values are being explained or predicted based on the explanatory variable.
- •The dependent variable whose internal variation is the main property we seek to explain.
Extrapolation
Predicting a response value using a value for the explanatory variable that is beyond the range of x-values used to create the regression model, resulting in less reliable predictions.
- •Highly suspect because we cannot confirm that the uniform rate of change persists beyond our sample parameters.
Least-squares regression line
A linear model that minimizes the sum of squared residuals to find the best-fitting line through a set of data points.
- •This path is mathematically locked into intersecting the precise group coordinate mean center (x̄, ȳ).
Linear regression model
An equation that uses an explanatory variable to predict a response variable in a linear relationship.
- •The core tool used to mathematically describe stable straight bivariate relationships.
Predicted value
The estimated response value obtained from a regression model, denoted as ŷ.
- •Represents an idealized model expectation along the trend line rather than a factual empirical tracking point.
Slope
The value b in the regression equation ŷ = a + bx, representing the rate of change in the predicted response for each unit increase in the explanatory variable.
- •Requires tentative framing like "predicted expansion" or "on average" to secure full AP context credit.
y-intercept
The value a in the regression equation ŷ = a + bx, representing the predicted response value when the explanatory variable equals zero.
- •Often maps to an illogical baseline condition if an input value of zero represents a physical or structural impossibility.
5.4 - Residuals
Key Terms & Definitions
Linear model
A mathematical representation of the linear relationship between two variables.
- •Employed to standardize, interpret, and calculate estimated behaviors across a linear domain space.
Predicted value
The estimated response value obtained from a regression model, denoted as ŷ.
- •Represents an idealized model expectation along the trend line rather than a factual empirical tracking point.
Actual value
The observed or measured response value in a dataset, denoted as y.
- •Serves as the concrete data factual baseline used to judge the relative accuracy of model trends.
Bivariate data
Data involving two variables, typically represented as ordered pairs (x, y) to examine the relationship between them.
- •Can link categorical groups (Unit 2) or explore numerical patterns via scatterplots (Unit 5).
Form of association
The pattern or type of relationship between two variables, such as linear, curved, or no relationship.
- •Identifying the form dictates whether fitting a simple linear equation is mathematically logical.
Randomness in residuals
The absence of a clear pattern in a residual plot, indicating that a linear model is appropriate for the data.
- •The absence of clear shapes or bends in error tracking confirms a steady, linear trend rate.
Residual
The difference between the actual observed value and the predicted value in a regression model, calculated as residual = y − ŷ.
- •Positive values reveal the line underpredicted reality; negative results index an overprediction trend.
Residual plot
A scatterplot of the residuals against the explanatory variable, used to assess whether a linear model is appropriate.
- •Any visible curvature or expansion pattern signals that standard linear formulas are mathematically flawed.
5.5 - Least-Squares Regression
Key Terms & Definitions
Explanatory variable
A variable whose values are used to explain or predict corresponding values for the response variable.
- •Its values are treated as the given conditions used to explain variance in the accompanying response metric.
Response variable
A variable whose values are being explained or predicted based on the explanatory variable.
- •The dependent variable whose internal variation is the main property we seek to explain.
Correlation
A numerical measure (r) that describes the strength and direction of a linear relationship between two variables, ranging from -1 to 1.
- •Sign tracks the direction, while absolute proximity to 1 validates the relative strength of the linear model.
Least-squares regression line
A linear model that minimizes the sum of squared residuals to find the best-fitting line through a set of data points.
- •This path is mathematically locked into intersecting the precise group coordinate mean center (x̄, ȳ).
Predicted value
The estimated response value obtained from a regression model, denoted as ŷ.
- •Represents an idealized model expectation along the trend line rather than a factual empirical tracking point.
Slope
The value b in the regression equation ŷ = a + bx, representing the rate of change in the predicted response for each unit increase in the explanatory variable.
- •Requires tentative framing like "predicted expansion" or "on average" to secure full AP context credit.
y-intercept
The value a in the regression equation ŷ = a + bx, representing the predicted response value when the explanatory variable equals zero.
- •Often maps to an illogical baseline condition if an input value of zero represents a physical or structural impossibility.
Residual
The difference between the actual observed value and the predicted value in a regression model, calculated as residual = y − ŷ.
- •Positive values reveal the line underpredicted reality; negative results index an overprediction trend.
Coefficient of determination
The value r², which represents the proportion of variation in the response variable that is explained by the explanatory variable in the regression model.
- •Interpreted uniformly via the script: "X% of the variation in response is accounted for by the linear relationship with input."
Coefficients
The numerical values in a regression equation that represent the slope and y-intercept of the least-squares regression line.
- •Extracted from software printouts to assemble the tracking model equation ŷ = a + bx.
Parameter
A numerical summary that describes a characteristic of an entire population.
- •Remains hidden in empirical practice; our sample coefficients act as provisional estimates for these metrics.
Sample standard deviation
The standard deviation calculated for a sample, denoted by s.
- •Used to manually compute line slopes when coupled with correlation metrics via b = r(s_y / s_x).
Simple linear regression
A regression model that describes the linear relationship between one explanatory variable and one response variable.
- •Restricted entirely to single variable pairings, separating it from advanced multi-variable tracking procedures.