Last updated July 07, 2026 · Reviewed by the AnvayaPrep team
Introduction
Data Analysis and Statistics is the unit covering the full scope of statistical reasoning tested in the SAT Problem Solving and Data Analysis domain, which comprises approximately 29% of SAT Math. The unit's 27 topics span measures of center and spread (mean, median, mode, standard deviation basics, interquartile range, outliers, box plots), data displays (histograms, dot plots, two-way tables, frequency tables, scatterplots), regression and association (line of best fit, regression interpretation, residuals basics, correlation), statistical inference (sampling, random sampling, bias, survey interpretation, margin of error, experimental design), and analytical reasoning (conditional percentages, causation, data comparison, data inference, SAT statistics traps).
Learning Objectives
- Calculate and interpret the mean, median, and mode of a dataset, and identify which measure is most appropriate in context
- Determine the effect of adding, removing, or changing data points on the mean and median
- Calculate and interpret standard deviation as a measure of spread, and understand that higher standard deviation indicates more variability
- Read and extract information from histograms, box plots, dot plots, two-way tables, and scatterplots
- Use a line of best fit to make predictions, and identify whether the model overestimates or underestimates using residuals
- Distinguish between positive correlation, negative correlation, and no correlation in scatterplots
- Evaluate whether a sampling method is likely to produce a representative sample or a biased sample
- Distinguish between correlation and causation: a study showing correlation does not prove that one variable causes the other
High-Yield Concepts
| Concept | What It Tests | Key Rule |
|---|---|---|
| Mean | Calculating the arithmetic average | Mean = (sum of all values) / (count of values) |
| Median | Finding the middle value | Order the data; median is the middle value (or average of two middle values for even count) |
| Effect of outliers on mean vs. median | Understanding which measure is resistant | Outliers pull the mean toward them but have little effect on the median |
| Standard deviation | Measuring spread or variability | Larger SD = more spread out; adding/removing values affects SD |
| Interquartile range | Measuring middle 50% spread | IQR = Q3 - Q1; resistant to outliers |
| Two-way tables | Reading conditional percents | Identify the correct total (row, column, or grand total) before computing percentages |
| Line of best fit | Making predictions from scatterplot data | Use the slope-intercept form to predict; residual = actual - predicted |
| Correlation vs. causation | Interpreting association claims | Correlation does not imply causation; only randomized controlled experiments can establish causation |
| Random sampling | Evaluating study validity | A random sample allows generalization to the population; a biased sample does not |
| Margin of error | Understanding survey precision | Results reported as "p% +/- m%" mean the true value is likely within m percentage points of p% |
Study Strategy
Begin with the three measures of center: mean (sum divided by count), median (middle value after ordering), and mode (most frequent value). The SAT asks you to calculate these, compare them, and analyze how changes to the dataset affect them. Outliers increase the mean significantly but barely move the median, making the median a better measure of center for skewed data.
Then study the displays: histograms (bars showing frequency distribution), box plots (five-number summary: min, Q1, median, Q3, max), dot plots (individual data points), and two-way tables (rows and columns showing two categorical variables). Know how to extract counts, percents, and conditional percents from each.
For scatterplots, practice reading the line of best fit (line of regression) to make predictions. The slope tells you how y changes as x increases by one unit. A residual is the difference between the actual y-value and the predicted y-value from the line: positive residual means the actual value is above the line, negative means below.
The most conceptually tested area is statistical inference. The key principle: a random sample can be used to generalize to the population it was drawn from. A biased (non-random) sample cannot. Correlation studies establish association but not causation; only a randomized controlled experiment can establish causation.
For conditional percentages from two-way tables, always identify the denominator (the total for the specific condition) before computing. "Of students who passed, what percent were female?" uses the number who passed as the denominator, not the total number of students.
Common Mistakes
Confusing mean and median when data is skewed. For a symmetric distribution, mean and median are close. For a right-skewed distribution (long tail to the right, such as income data), the mean is higher than the median because extreme high values pull the mean up. The SAT tests this by asking which measure better represents a dataset or by asking what happens to the mean vs. median when an outlier is added.
Using the wrong total in two-way table percent problems. Two-way table questions specify which group to focus on. "What percent of all students passed?" uses the grand total. "What percent of female students passed?" uses the total number of female students. Using the wrong total is the most common error on two-way table questions.
Confusing residuals. A residual is actual minus predicted. If the actual value is above the line of best fit, the residual is positive; the model underestimated. If the actual value is below the line, the residual is negative; the model overestimated. Students often reverse these.
Claiming causation from correlation. If a study finds that students who eat breakfast score higher on tests, this does not prove that eating breakfast causes higher scores. A third variable (such as overall health habits) might explain both. Only a randomized controlled experiment -- where participants are randomly assigned to groups -- can establish causation.
Generalizing from a biased sample. If a survey about school lunch preferences was given only to students eating in the cafeteria, the results cannot be generalized to all students. The SAT presents non-random samples and asks whether the conclusions are valid; the answer is always that you can only generalize to the population the sample was drawn from.
Exam Tips
On mean problems where the sum is more useful than the mean, use the relationship: sum = mean x count. If the mean of 8 values is 12, the sum is 96. If a 9th value is added and the new mean is 13, the new sum is 117, so the added value is 117 - 96 = 21. This algebraic approach to mean problems is faster than repeatedly recalculating means.
When a SAT question mentions a study and asks what conclusion is supported, watch for two traps: (1) generalization to the wrong population (a study of gym members cannot generalize to all adults), and (2) causation language (a correlation study never supports a "causes" conclusion, only an "association" conclusion). Both are standard SAT wrong-answer constructions.
The five-number summary for a box plot: Minimum, Q1 (25th percentile), Median (50th percentile), Q3 (75th percentile), Maximum. The IQR = Q3 - Q1. An outlier is often defined as any value more than 1.5 * IQR below Q1 or above Q3. A longer whisker or box on one side indicates the data is skewed in that direction.
Standard deviation measures how spread out data is from the mean. You do not need to calculate it from scratch on the SAT, but you must understand: datasets with values clustered near the mean have small SD, datasets with values spread far from the mean have large SD. If two datasets have the same mean, the one with more variation has higher SD. Adding a constant to every value does not change SD; multiplying every value by a constant multiplies SD by that constant.
Sign up free to keep reading
Create a free AnvayaPrep account to finish this SAT guide on Data Analysis and Statistics — plus flashcards and practice questions.