Digital SAT Math · Problem-Solving & Data Analysis

Two-variable data (scatterplots & models)

A scatterplot puts two numbers on every subject — a tree’s age and its harvest, a stop’s length and its checkouts — and then the test draws a line through the cloud and asks about the line. That is the whole trick of this skill: the dots are what happened, the model is what was expected, and almost every question here turns on knowing which of the two the stem just asked for. Get that straight and the arithmetic is one substitution or one subtraction.

On the test

DomainProblem-Solving and Data Analysis (score report)
CB skillTwo-variable data: models and scatterplots
What it looks likeA scatterplot with a line of best fit drawn, or a short table plus a model equation, and one question about the model
Often asked“What does the number … represent?”, “What value does the model predict…?”, “How much greater than predicted…?”, “Which model best fits these data?”
FormatMultiple choice and student-produced response
CalculatorAllowed throughout; genuinely useful here, because the models carry decimals

Recognition cues: scatterplot, line of best fit, the model predicts, residual, actual, estimated by the line, best fits the data, for each additional.

Pattern recognition

Two objects live on the screen, and every question belongs to exactly one of them.

  • The dots are the observed data. Each dot is one subject with its two measurements. A dot’s yy is an actual value.
  • The line (or curve) is the model. It exists at every xx, including ones nobody measured. A value read off the line is a predicted or estimated value.

Circle the word in the stem that decides which object you need. Recorded, observed, actual, measured send you to a dot. Predicted, estimated, the model, the line of best fit send you to the line — even at an xx where a dot happens to sit.

Method

  1. Label both axes with units. Write them down: ”xx = age in years, yy = harvest in kilograms”. Every later sentence you produce has to survive being read with those units attached, which kills half of the wrong answers before you look at them.
  2. Say the slope in context. Slope is the change in yy for a one-unit increase in xx:
slope=change in ychange in x⇒"for each additional 1 x-unit, the model predicts m more y-units."\text{slope} = \frac{\text{change in }y}{\text{change in }x} \quad\Rightarrow\quad \text{"for each additional 1 } x\text{-unit, the model predicts } m \text{ more } y\text{-units."}

The intercept is the model’s value at x=0x = 0 — real only if x=0x = 0 is a sensible thing for the context. 3. Predict off the line, read actuals off the dots. To predict, substitute xx into the equation. If the picture gives the line but not the equation, use two lattice points on the line — never two scattered dots, which would give you a slope the model does not have. 4. Residual = actual − predicted, and the sign is the answer.

residual=yactual−ypredicted\text{residual} = y_{\text{actual}} - y_{\text{predicted}}

Positive means the dot sits above the line and the model underestimated. Negative means the dot sits below and the model overestimated. Check your sign against the picture every single time. 5. Choose the model from the shape, before the numbers. A straight run of dots is linear; a run that bends steeply upward (or flattens toward an asymptote) is exponential. Confirm it with the data: constant differences → linear, constant ratios → exponential.

Words in the stemWhat to do
what does … representone sentence, with units, about a one-unit increase in xx
predicted / estimated by the modelsubstitute xx into the equation
actual / recorded / observedread the dot, not the line
residualactual − predicted, keep the sign
how many more than predictedactual − predicted, and expect a positive answer
overestimatespredicted > actual — the dot is below the line
best fits the datadifferences constant → linear; ratios constant → exponential

Worked example 1 — reading a model, then predicting from it

Stem. A kiosk recorded, for each of 12 summer days, the day’s high temperature tt (in degrees Fahrenheit) and the number of scoops of gelato sold ss. The high temperatures ranged from 74 °F to 96 °F. The line of best fit for the data is

s=14t−560.s = 14t - 560.

(a) What does 14 represent? (b) How many scoops does the model predict for a day with a high of 88 °F?

Step 1 — units on both axes. tt is in degrees Fahrenheit; ss is in scoops. So the slope’s units are scoops per degree.

Step 2 — (a) say the slope in context. A one-degree increase in the high temperature goes with a predicted increase of 14 scoops sold. Notice what the sentence is not: it is not a claim about any particular day, and it is not “14 scoops were sold”.

Step 3 — (b) substitute.

s=14(88)−560=1232−560=672s = 14(88) - 560 = 1232 - 560 = 672

Check. At 74 °F the model gives 14(74)−560=47614(74) - 560 = 476 and at 96 °F it gives 784784, so 672 at 88 °F sits sensibly between them and nearer the top, which is where 88 sits in the temperature range. Answer: (a) each additional degree goes with 14 more scoops predicted; (b) 672 scoops.

Trap watch. The intercept −560-560 is not “the kiosk sells −560-560 scoops at 0 °F” in any meaningful sense — t=0t = 0 is nowhere near the plotted data, and negative scoops are not a thing. An intercept is only interpretable when x=0x = 0 is inside, or at least adjacent to, the range you actually measured. The other trap is inverting the slope: “for each additional scoop sold, the temperature rises 14 degrees” reverses the axes and produces a sentence with the wrong units.

Worked example 2 — a residual, with its sign

Stem. A cycling coach plotted, for each of 14 riders, the number of interval workouts completed ww and the improvement pp, in watts, in the rider’s 20-minute power. The line of best fit is

p=2.6w+11.p = 2.6w + 11.

One rider completed 15 interval workouts and improved by 44 watts. What is the residual for that rider?

Step 1 — predicted, from the line.

ppredicted=2.6(15)+11=39+11=50 wattsp_{\text{predicted}} = 2.6(15) + 11 = 39 + 11 = 50 \text{ watts}

Step 2 — actual, from the rider. The stem hands it to you: 44 watts. That is the dot.

Step 3 — subtract in the required order.

residual=44−50=−6 watts\text{residual} = 44 - 50 = -6 \text{ watts}

Check. The residual is negative, so the dot should sit below the line — and it does: this rider improved by 44 watts where the model expected 50. The model overestimated by 6 watts, which is the same fact stated with the sign moved into the word “overestimated”. Answer: −6-6 watts.

Trap watch. +6+6 is predicted minus actual, the reversed subtraction, and it will be sitting in the choice list. So will 50, which answers “what did the model predict” — a question nobody asked. And so will 44, which is just the number already printed in the stem. If your answer to a residual question is a number you copied rather than computed, re-read the question.

Worked example 3 — choosing the model from the shape

Stem. Two seed libraries tracked the number of items in their lending collections at the end of each of 5 months.

  • Library M: 1200, 1450, 1700, 1950, 2200
  • Library N: 1024, 1280, 1600, 2000, 2500

Which type of model fits each, and what does each model predict for month 6?

Step 1 — first differences. For M: 250,250,250,250250, 250, 250, 250 — constant. For N: 256,320,400,500256, 320, 400, 500 — growing, so N is not linear.

Step 2 — ratios. For N: 12801024=1.25\frac{1280}{1024} = 1.25, 16001280=1.25\frac{1600}{1280} = 1.25, 20001600=1.25\frac{2000}{1600} = 1.25, 25002000=1.25\frac{2500}{2000} = 1.25 — constant. A constant ratio is exactly what an exponential model means.

Step 3 — predict each one in its own currency. M adds the same amount again: 2200+250=24502200 + 250 = 2450. N multiplies by the same factor again: 2500×1.25=31252500 \times 1.25 = 3125.

Check. Both collections grew over the five months, and both grew by a lot; only one of them grew by a constant amount. M gained 1000 items across four steps and N gained 1476 across the same four, and N’s steps got bigger every time — the signature of a curve, not a line. Answer: M is linear and predicts 2450 items; N is exponential and predicts 3125 items.

Trap watch. “It goes up every month, so it’s linear” is the error the whole example exists to kill — increasing is not the same as increasing at a constant rate. The second trap is doing the right diagnosis and then the wrong arithmetic: adding N’s most recent difference, 2500+500=30002500 + 500 = 3000, instead of multiplying. Once you have decided the model is exponential, the step is a multiplication, and it must be by the ratio you actually measured — 1.25, not 2.

Practice

Answer before you open the explanation. Two items are student-produced response (type the number, no choices), matching the real test, and five are built on a display — four scatterplots and one table — because that is how the test presents this skill; one reviewed competitor’s item set on this exact CB skill was figure-based in every single stem. Every wrong choice below is a specific error with a name.

12 questions — 10 multiple choice, 2 student-produced response. Every wrong choice has its own explanation.

Question 1 Warm-up

A city recorded, for each of its kerbside recycling routes, the number of households h on the route and the total mass m, in kilograms, of recycling collected from that route in one week. The line of best fit for the data is m = 0.62h + 14. Which of the following is the best interpretation of the number 0.62 in this context?

Show the answer Choice C

Why it is right

In m = 0.62h + 14 the variable h counts households and m is measured in kilograms, so the slope 0.62 carries the units kilograms per household. A slope is the change in the output for a one-unit increase in the input, so 0.62 says that a route with one more household is predicted to yield 0.62 kilograms more recycling in a week. That is a statement about the model's rate of change, not about any single route that was measured.

Why each other choice fails

Choice A
That describes the intercept 14, the model's output when h = 0, not the slope. The two constants in a linear model answer different questions: 14 is a mass in kilograms, while 0.62 is a mass per household.
Choice B
This inverts the axes. Households per kilogram would be the slope of a model with the variables swapped; here the input is households and the output is kilograms, so the units of 0.62 are kilograms per household, not the reverse.
Choice D
That defines a residual, which is a property of one particular route and its data point. The slope is a feature of the fitted line itself and says nothing about how far any individual route sits above or below it.

Question 2 Standard

The scatterplot shows, for each of 10 trees in an orchard, the age a of the tree in years and the mass m, in kilograms, of apples harvested from that tree last season. The line of best fit shown has equation m = 3a + 8. According to the line of best fit, what mass of apples, in kilograms, is predicted for a tree that is 12 years old?

0 2 4 6 8 10 12 14 0 10 20 30 40 50 m = 3a + 8 a = 12 Age of tree (years) Apples harvested (kg)
Age and harvest for ten trees in the orchard, with the line of best fit and its value at 12 years.
Show the answer Choice B

Why it is right

A prediction comes from the model, so substitute the age into the equation of the line of best fit rather than looking for a nearby dot. With a = 12, m = 3(12) + 8 = 36 + 8 = 44 kilograms. The open marker on the line at a = 12 sits at 44, which confirms the substitution against the picture; no tree of that age was plotted, and the model still supplies a value for one.

Why each other choice fails

Choice A
36 is 3 times 12 with the intercept dropped. The model is m = 3a + 8, so the constant 8 must be added after multiplying; leaving it out understates every prediction the model makes by exactly 8 kilograms.
Choice C
40 kilograms is the harvest actually recorded for the oldest tree plotted, the 11-year-old. That is a data point, not the model's prediction, and the question asked what the line of best fit predicts for a 12-year-old tree.
Choice D
41 is the model's prediction at a = 11, not a = 12: 3(11) + 8 = 41. Substituting the largest age that appears in the scatterplot instead of the age the question names shifts the answer by one year's worth of growth.

Question 3 Standard

The table gives the number of hours h that a food stall stayed open and the number of meals n it served on each of five days. The line of best fit for these data is n = 26h + 15. On which day did the stall serve the most meals relative to the number predicted by the model for that day?

Hours openMeals served
4125
5140
6180
7190
8225
Hours open and meals served on each of five days.
Show the answer Choice D

Why it is right

"Relative to the number predicted" means compare each day's actual count with the model's value at that day's hours, that is, compute actual minus predicted for every row. The predictions are 26(4) + 15 = 119, 26(5) + 15 = 145, 26(6) + 15 = 171, 26(7) + 15 = 197 and 26(8) + 15 = 223. Subtracting each from the actual count gives 125 - 119 = 6, 140 - 145 = -5, 180 - 171 = 9, 190 - 197 = -7 and 225 - 223 = 2. The largest of these is 9, on the 6-hour day.

Why each other choice fails

Choice A
The 4-hour day has the highest meals-per-hour rate, 125 / 4 = 31.25 against 30 on the 6-hour day, but a rate is not a comparison with the model. Measured against the line of best fit, that day beat its prediction by only 6 meals.
Choice B
The 7-hour day is the day the stall did worst against the model, not best: its prediction of 197 exceeds the 190 meals actually served, giving a residual of -7. Choosing it means subtracting predicted minus actual and then reading the largest result as the best day.
Choice C
The 8-hour day served the most meals in absolute terms, 225, but the model already expects a big number when the stall stays open longest. Its prediction of 223 was beaten by only 2 meals, the smallest positive gap in the table.

Question 4 Standard

A shop recorded, for each of 24 weeks, the shelf price p, in dollars, of a jar of jam and the number of jars s that were sold that week. The line of best fit for the data is s = -18p + 250. Which of the following is the best interpretation of the number -18 in this context?

Show the answer Choice A

Why it is right

The input p is a price in dollars and the output s is a number of jars, so the slope -18 has units of jars per dollar. A slope is the predicted change in the output when the input rises by one unit, and the negative sign means the change is a decrease: raising the shelf price by $1 goes with 18 fewer jars sold per week according to the model. Nothing here is a claim about any individual week that was recorded.

Why each other choice fails

Choice B
This reads the slope with the axes swapped. Dollars per jar would be the slope of a model that predicted price from sales; in s = -18p + 250 the price is the input, so -18 must be read as jars per dollar.
Choice C
The model's value at p = 0 is the intercept, s = -18(0) + 250 = 250 jars, not 18. Confusing the slope with the intercept swaps a rate of change for a single predicted output.
Choice D
The slope describes the model's rate of change per dollar, not the gap between two particular weeks in the data. Two recorded weeks differ by their actual sales and by however many dollars separate their prices, which -18 alone does not determine.

Question 5 Standard

The scatterplot shows, for each of 9 bookmobile stops, the length of the stop in minutes and the number of books checked out during that stop. The line of best fit shown has equation y = 1.5x + 10, where x is the length of the stop in minutes. At the stop that lasted 40 minutes, how many more books were checked out than the number predicted by the line of best fit?

0 10 20 30 40 50 0 10 20 30 40 50 60 70 80 90 y = 1.5x + 10 40-minute stop Length of stop (minutes) Books checked out
Length of stop and books checked out at nine bookmobile stops, with the line of best fit and the gap at the 40-minute stop marked.
Show the answer Choice A

Why it is right

The question asks for a gap, so both numbers are needed. The predicted value comes from the model: at x = 40, y = 1.5(40) + 10 = 60 + 10 = 70 books. The actual value comes from the dot plotted at 40 minutes, which sits at 78 books. The difference is 78 - 70 = 8, and the dashed segment on the plot is exactly that gap. Because the dot lies above the line, the answer is positive: the model underestimated this stop by 8 books.

Why each other choice fails

Choice B
18 comes from predicting with 1.5(40) = 60 and forgetting the intercept, then computing 78 - 60. Dropping the constant term makes every prediction 10 books too low and inflates every gap by the same 10.
Choice C
70 is the predicted number of books, not the amount by which the actual value exceeds it. It answers "what does the model say" rather than "how much more than the model", so the final subtraction is missing.
Choice D
78 is the actual number of books checked out at that stop, read straight off the dot. It is one of the two numbers the subtraction needs, not the answer, and it ignores the model entirely.

Question 6 Harder

The scatterplot shows, for each of 8 rowing crews, the number of practice sessions p a crew attended and the crew's technique score s. The line of best fit shown has equation s = 4p + 6. One crew is highlighted. What is the residual for the highlighted crew, where the residual is the actual score minus the score predicted by the line of best fit?

0 1 2 3 4 5 6 7 8 9 0 5 10 15 20 25 30 35 40 45 s = 4p + 6 highlighted crew Practice sessions attended Technique score
Practice sessions and technique scores for eight crews, with the line of best fit and the highlighted crew's gap to the line.
Show the answer Choice B

Why it is right

The highlighted crew attended 7 practice sessions and scored 30, so the actual value is 30. The model's value at that input is s = 4(7) + 6 = 28 + 6 = 34. A residual is actual minus predicted, in that order, so the residual is 30 - 34 = -4. The sign matches the picture: the highlighted dot sits below the line, which is what a negative residual means, and it says the model overestimated this crew by 4 points.

Why each other choice fails

Choice A
4 is predicted minus actual, 34 - 30, the subtraction taken in the wrong order. It has the right magnitude but claims the crew scored above the line, which contradicts the dot's position below it.
Choice C
30 is the crew's actual score, one of the two inputs to the subtraction rather than its result. A residual measures a distance from the line, so an answer copied straight out of the data cannot be one.
Choice D
34 is the score the model predicts for a crew that attended 7 sessions. It answers what the line says, not how far the actual score fell from it, so the final subtraction has been skipped.

Question 7 Harder Student-produced response

A utility recorded, for each of 40 rooftop solar arrays, the number of panels n in the array and the electricity E, in kilowatt-hours, that the array generated on one clear day. The line of best fit for the data is E = 1.8n + 9. According to this model, how many panels would an array need in order to generate 99 kilowatt-hours on such a day?

Show the answer 50

Why it is right

This is a prediction run backwards: the output is given and the input is wanted, so set the model equal to the target and solve. From 1.8n + 9 = 99, subtract the intercept first to get 1.8n = 90, then divide by the slope to get n = 90 / 1.8 = 50 panels. Checking forwards confirms it: 1.8(50) + 9 = 90 + 9 = 99 kilowatt-hours, exactly the target the question named.

Answers students type instead

55
Divides 99 by the slope without removing the intercept first: 99 / 1.8 = 55. That answers the question for the model E = 1.8n, which has no constant term, and it overstates the panel count because 9 of the 99 kilowatt-hours never came from the slope.
60
Adds the intercept instead of subtracting it, giving (99 + 9) / 1.8 = 60. The equation is 1.8n + 9 = 99, so undoing the addition of 9 means subtracting it from both sides, not adding it.
90
Stops at 1.8n = 90 and reports 90. That is the number of kilowatt-hours attributable to the panels, not the number of panels; dividing by the slope 1.8 is the step that converts it into a count.

Question 8 Harder

During each of the first four weeks after a compost heap was seeded, a gardener estimated the worm population and recorded 400, 600, 900, and 1350 worms. Which of the following best describes a model for these data, and why?

Show the answer Choice D

Why it is right

Test differences first, then ratios. The week-to-week differences are 600 - 400 = 200, 900 - 600 = 300 and 1350 - 900 = 450, which are not constant, so no linear model fits. The ratios are 600 / 400 = 1.5, 900 / 600 = 1.5 and 1350 / 900 = 1.5, all equal. A constant ratio between successive equally spaced values is exactly what an exponential model means, so the population is best modelled by an exponential function with growth factor 1.5.

Why each other choice fails

Choice A
The reason is false for these data. The increases are 200, then 300, then 450, so the population does not rise by the same amount each week, and a linear model would have to fit a constant difference that is not there.
Choice B
Increasing is not the same as increasing linearly. Every model that grows increases each week, including this exponential one, so "it goes up" cannot distinguish a line from a curve; only the pattern of differences and ratios can.
Choice C
The family is right and the reason is not. A constant amount added each period describes a linear model; what makes this data set exponential is a constant multiplier, 1.5, applied each period.

Question 9 Harder

The scatterplot shows the area, in hectares, of each of 11 restored wetland ponds and the number of frog species recorded at that pond. The line of best fit shown has equation y = 2x + 3, where x is the pond's area in hectares. For how many of the 11 ponds does the line of best fit overestimate the number of frog species recorded?

0 1 2 3 4 5 6 7 8 9 10 11 12 0 5 10 15 20 25 30 y = 2x + 3 Pond area (hectares) Frog species recorded
Area and frog species count for eleven restored wetland ponds, with the line of best fit.
Show the answer Choice C

Why it is right

The model overestimates a pond when its predicted value is greater than the recorded value, which is exactly when the dot lies below the line. Comparing each dot with 2x + 3 gives predictions 5, 7, 9, 11, 13, 15, 17, 19, 21, 23 and 25 against recorded counts 6, 6, 10, 10, 14, 14, 18, 18, 22, 21 and 26. The prediction is larger at areas of 2, 4, 6, 8 and 10 hectares, which is 5 ponds; at the other 6 the recorded count is larger.

Why each other choice fails

Choice A
1 counts only the pond at 10 hectares, whose dot sits two species below the line and is the one obviously wide gap. The five dots that sit a single species below the line are overestimated too; being close to the line is not the same as being on it.
Choice B
6 is the number of ponds the model underestimates, the dots lying above the line. Overestimate means the model's number is the larger one, so the dots that count are the ones below the line, and there are 5 of those.
Choice D
11 treats every pond as overestimated because no dot lies exactly on the line. The line supplies a prediction for all 11 ponds, but it is above the recorded value for only 5 of them and below it for the other 6.

Question 10 Hardest

A researcher plotted, for each of 22 glacial lakes, the lake's maximum depth d, in metres, against its July surface temperature T, in degrees Celsius. The maximum depths of the lakes plotted range from 5 metres to 60 metres, and the line of best fit for the data is T = -0.08d + 14.6. Which of the following statements is best supported by this model?

Show the answer Choice A

Why it is right

Substituting a depth inside the range that was actually measured gives a prediction the data supports: T = -0.08(40) + 14.6 = -3.2 + 14.6 = 11.4 degrees Celsius, and 40 metres sits comfortably between the 5-metre and 60-metre extremes of the plotted lakes. The word "predicted" is what makes the statement defensible: it describes the model's value for a lake of that depth, not a guarantee about any particular lake.

Why each other choice fails

Choice B
The arithmetic is the same as the key, but "every lake has" turns a prediction into a certainty. The dots scatter around the line rather than sitting on it, which is direct evidence that individual lakes of the same depth differ from the model's value.
Choice C
The substitution is carried out correctly, -0.08(300) + 14.6 = -9.4, but 300 metres is five times the deepest lake in the data. Extrapolating that far past the plotted range is unsupported, and here it produces a below-freezing surface temperature in July.
Choice D
A scatterplot and its fit line show that two quantities move together; they cannot establish that one produces the other. Deep lakes may be cooler for reasons the study never varied, and no lake's depth was changed to see what happened to its temperature.

Question 11 Hardest Student-produced response

A creamery models the moisture content M, as a percent, of a wheel of cheese from its age w, in weeks, using the line of best fit M = -0.75w + 46. One wheel that was 16 weeks old was measured at 33.4 percent moisture. What is the residual for that wheel, where the residual is the actual moisture content minus the moisture content predicted by the model?

Show the answer -0.6

Why it is right

Find the prediction first: at w = 16 the model gives M = -0.75(16) + 46 = -12 + 46 = 34 percent. The actual measurement is 33.4 percent. A residual is actual minus predicted, so the residual is 33.4 - 34 = -0.6 percentage points. The negative sign is part of the answer and is the informative part: this wheel was drier than the model expected, so the model overestimated its moisture by 0.6 of a percentage point.

Answers students type instead

34
The moisture content the model predicts for a 16-week-old wheel. It is one of the two numbers the residual needs, not the residual itself, so the final subtraction is missing.
0.6
Predicted minus actual, 34 - 33.4, the subtraction in the wrong order. The magnitude is right but the sign claims the wheel was wetter than predicted, which reverses what the measurement shows.
45.4
Uses -0.75(16) = -12 as the prediction, dropping the intercept 46, and then computes 33.4 - (-12) = 45.4. A residual is never larger than the values being compared; a result bigger than either input signals that a term was lost.

Question 12 Hardest

A camera network recorded the number of confirmed beaver sightings in a river valley in each of 6 consecutive years: 64, 96, 144, 216, 324, and 486. A scatterplot of year number against sightings curves upward. Which model type best fits these data, and what does that model predict for the seventh year?

Show the answer Choice B

Why it is right

The year-to-year differences are 32, 48, 72, 108 and 162, which grow rather than staying constant, so no line fits. The ratios are 96 / 64 = 1.5, 144 / 96 = 1.5, 216 / 144 = 1.5, 324 / 216 = 1.5 and 486 / 324 = 1.5, all equal, so the data are exponential with growth factor 1.5. Extending the pattern one more year multiplies rather than adds: 486 x 1.5 = 729 sightings in the seventh year.

Why each other choice fails

Choice A
648 is 486 + 162, adding the most recent difference again as though the counts rose by a fixed amount. The differences are 32, 48, 72, 108 and 162, so no single constant increase describes the data and the linear diagnosis fails before the arithmetic starts.
Choice C
The exponential diagnosis is right but the growth factor is taken as 2, giving 486 x 2 = 972. The factor has to be measured from the data, and every consecutive pair here has a ratio of 1.5, not 2.
Choice D
518 is 486 + 32, treating the very first difference as if it held for every year. That underestimates badly precisely because the differences are growing, and it again names a linear model for a data set with no constant difference.

Common mistakes

  1. Reading a dot when the stem said “predicted” — the model’s value at xx comes from substituting into the equation, even when a data point happens to sit at that same xx. The dot is what happened; the line is what was expected.
  2. Residual computed backwards — residual is actual minus predicted, in that order. The reversed subtraction gives the right magnitude with the wrong sign, and that value is always among the choices.
  3. “Overestimates” read as “the dot is above the line” — an overestimating model predicted too much, so its prediction is above and the dot is below, and the residual is negative.
  4. Slope stated with the units inverted — “for each additional kilogram, 0.62 more households” reverses the axes. Say yy-units per one xx-unit and the sentence either survives or collapses.
  5. Slope computed from two scattered dots — two data points define a line through those two points, not the line of best fit. If you must measure the slope from a picture, use two points that lie on the drawn line.
  6. The intercept interpreted where x=0x = 0 is nonsense — a negative predicted count, or a temperature at zero panels, is a sign that the intercept is a fitting constant here, not a meaningful starting value.
  7. Increasing mistaken for linear — every data set on this page increases; only some increase by a constant amount. Check first differences, then ratios, before naming the model.
  8. The growth factor guessed instead of measured — after correctly deciding a model is exponential, students often double. Compute a ratio from consecutive values and use that number.
  9. Extrapolating far outside the plotted range as if it were certain — the equation returns a value at any xx; the evidence for it stops where the dots stop.
  10. Treating the model as a rule about individuals, or as a cause — a fit line describes an average tendency across the whole cloud. It does not promise any one subject’s value, and correlation in a scatterplot is never on its own evidence that one variable causes the other.

FAQ

Do I ever have to find the line of best fit myself? Not on this skill. The Digital SAT draws the line or prints its equation and asks what it means or what it predicts. Fitting a line to points — from two coordinates, from a graph, from a table — is a linear-functions task, and Desmos will do a regression for you if a question ever genuinely calls for one.

What exactly is a residual, in one sentence? The vertical distance from a data point to the line of best fit, signed: actual value minus predicted value. Positive means the point is above the line, negative means below. A residual of zero means the point sits exactly on the line, which the test occasionally builds a question around.

How do I tell an exponential scatter from a linear one when the picture is small? Ignore the picture and use the numbers. Take consecutive yy-values at equally spaced xx-values and subtract: if the differences are roughly equal, it is linear. Then divide instead: if the ratios are roughly equal, it is exponential. On a plot, the visual cue is that an exponential run’s steps get visibly steeper as you move right, while a linear run’s steps stay the same height.

Can I type a negative number into a grid-in? Yes — the student-produced response field accepts a minus sign, which matters here because a residual can legitimately be negative. Type the sign; do not report the absolute value unless the question asked for a distance or for “how much greater”.

The question asks how much greater the actual value is than the predicted one. Is that a residual? It is the same subtraction, and if the wording is “how much greater than predicted”, the expected answer is positive. If it is phrased as “the residual”, give the signed value. Reading which of the two is wanted takes three seconds and is worth more than any arithmetic on the page.