Appendix E — Exercise Solutions

Solutions to the odd-numbered end-of-chapter exercises are provided below, organized by chapter. Even-numbered exercises are reserved for graded assignments; full solutions are available to instructors.

E.1 Chapter 1 — Introduction to Data

Solutions to the odd-numbered exercises in Chapter 1 — Introduction to Data.

  1. 23 observations and 7 variables.
  2. (a) “Is there an association between air pollution exposure and preterm births?” (b) 143,196 births in Southern California between 1989 and 1993. (c) Concentrations of carbon monoxide, nitrogen dioxide, ozone, and particulate matter with an aerodynamic diameter of 10 micrometres or less (PM\(_{10}\)) measured at air quality monitoring stations, as well as length of gestation. Continuous numerical variables.
  3. (a) “What is the effect of gamification on learning outcomes compared to traditional teaching methods?” (b) 365 college students taking a statistics course. (c) Gender (categorical), level of studies (categorical, ordinal), academic major (categorical), expertise in English language (categorical, ordinal), use of personal computers and games (categorical, ordinal), treatment group (categorical), score (numerical, discrete).
  4. (a) Treatment: \(10/43 = 0.23 \rightarrow 23\%\). (b) Control: \(2/46 = 0.04 \rightarrow 4\%\). (c) A higher percentage of patients in the treatment group were pain free 24 hours after receiving acupuncture. (d) It is possible that the observed difference between the two group percentages is due to chance — that is, sampling variability alone could produce a gap this large even if acupuncture had no real effect. (e) Explanatory: whether the patient received acupuncture. Response: whether the patient was pain free after 24 hours.
  5. (a) Experiment; researchers manipulated whether a daycare center imposed a fine and observed how parents’ late-pickup behavior changed. (b) 200 cases: the weekly observations of the 10 daycare centers over the 20 weeks. (c) Number of late pickups (numerical, discrete). (d) Week (numerical, discrete), group (categorical, nominal — treatment or control), and study period (categorical, ordinal — before, during, after).
  6. (a) 344 cases (penguins) are included in the data. (b) There are 4 numerical variables: bill length, bill depth, and flipper length (measured in millimeters) and body mass (measured in grams). They are all continuous. (c) There are 3 categorical variables: species (Adelie, Chinstrap, Gentoo), island (Torgersen, Biscoe, Dream), and sex (female, male).
  7. (a) Airport ownership status (public/private), airport usage status (public/private), region (Central, Eastern, Great Lakes, New England, Northwest Mountain, Southern, Southwest, Western Pacific), latitude, and longitude. (b) Airport ownership status: categorical, not ordinal. Airport usage status: categorical, not ordinal. Region: categorical, not ordinal. Latitude: numerical, continuous. Longitude: numerical, continuous.
  8. (a) Year, number of baby girls named Fiona born in that year, and nation. (b) Year (numerical, discrete), number of baby girls named Fiona born in that year (numerical, discrete), and nation (categorical, nominal).
  9. (a) County, state, driver’s race, whether the car was searched or not, and whether the driver was arrested or not. (b) All categorical, non-ordinal. (c) Response: whether the car was searched. Explanatory: race of the driver.
  10. (a) Observational study — researchers observed pet-name popularity without assigning any treatment. (b) Dog: Lucy. Cat: Luna. (c) Oliver and Lily. (d) Positive: as the popularity of a name among dogs increases, its popularity among cats tends to increase as well.

E.2 Chapter 2 — Data Collection and Study Design

Solutions to the odd-numbered exercises in Chapter 2 — Data Collection and Study Design.

  1. (a) Population mean, \(\mu_{2007} = \$52\); sample mean, \(\bar{x}_{2008} = \$58\). (b) Population mean, \(\mu_{2001} = 3.37\); sample mean, \(\bar{x}_{2012} = 3.59\).
  2. (a) Population: all births in Southern California. Sample: the 143,196 births recorded between 1989 and 1993. (b) If births in this time window at these locations are representative of all Southern California births, the results generalize to that population. Because the study is observational, it cannot establish a causal link between air-pollutant exposure and preterm birth.
  3. (a) The population of interest is all college students studying statistics; the sample is these 365 students. (b) The students were not randomly sampled from the broader population and come from two specific majors, so generalizing beyond similar students is a stretch. Because the study is an experiment (random assignment to gamification conditions), a causal statement about gamification’s effect within the studied group is justified.
  4. (a) Observation (a single case). (b) Variable. (c) Sample statistic — a sample mean, \(\bar{x}\). (d) Population parameter — a population mean, \(\mu\).
  5. (a) Observational. (b) Stratified sampling by section: randomly draw, say, 10 students from each of the four sections for a total sample of 40. This guarantees representation from every section.
  6. (a) Positive, non-linear, moderately strong: countries with higher internet-access rates tend to have higher life expectancy, but the increase levels off around 80 years. (b) Observational. (c) National wealth is a plausible confounder — richer countries can afford both widespread internet access and better health care. (Other reasonable confounders are acceptable.)
  7. (a) Simple random sampling is a reasonable choice — it very rarely goes wrong. (b) Stratifying by field of study is reasonable, since opinions may differ by major. (c) Clustering by age is a poor choice: students of similar ages likely hold similar opinions, so clusters would not be internally diverse (and clusters could be very unequal in size).
  8. (a) The cases are the 200 Mechanical Turk workers who took the survey. (b) Response: attitude toward the fictional microwave oven. (c) Explanatory: dispositional attitude. (d) The 200 participants were recruited randomly online. (e) Observational — participants were not randomly assigned to conditions. (f) No, no causal claim: the study is observational. (g) Generalization to Mechanical Turk workers is reasonable; broader generalization is not.
  9. (a) Simple random sample. Risk: non-response bias — students who reply may hold stronger opinions than those who don’t. (b) Convenience sample. Risk: undercoverage bias — friends are unlikely to represent the whole student body; non-response is also possible. (c) Convenience sample — same concerns as (b). (d) Multi-stage sampling (a random sample of classes, then all students in those classes). If the sampled classes are similar to the population of classes, the main remaining concern is non-response.
  10. (a) Response: exam performance. (b) Explanatory (treatment): light level, with three levels (fluorescent overhead, yellow overhead, no overhead / only desk lamps). (c) Whether the student wears glasses is used as a blocking variable.
  11. (a) Experiment. (b) Two explanatory (treatment) variables: light level (three levels) and noise level (no noise, construction noise, human chatter). (c) Because the researchers want equal representation of glasses-wearers and non-wearers, “wears glasses” is a blocking variable.
  12. Randomize and blind. One workable design: for each participant, prepare two identical cups — one regular Coke, one Diet Coke — labeled A and B, with the A/B \(\leftrightarrow\) Coke/Diet assignment randomized independently for each trial. Present the two cups in random order and ask the participant to rate each on how much they liked it. Neither the participant nor the person handing out the cups should know which beverage is which (double-blind). (Other reasonable designs accepted.)
  13. (a) Experiment. (b) Treatment: 25 g of chia seeds twice a day; control: an indistinguishable placebo. (c) Yes — blocking by gender. (d) Single-blind (patients were blinded to their treatment). (e) Because the study is a randomized experiment, causal statements about the studied participants are justified; however, the sample is not random, so the causal claim does not generalize to the broader population.
  14. (a) Non-response bias. The parents who returned the survey are likely those with easier schedules or fewer conflicts; the 85% figure understates the population’s difficulty. The school should follow up with non-responders (or take a small random sample of the entire class and pursue high response). (b) The 567 women reached three years later are almost certainly not representative — they are more likely to be homeowners, of higher socioeconomic status, etc. The researcher should compare respondents to the original sample and report the potential bias, not treat the follow-up as a random sample. (c) The study has no comparison group and is observational. People who run may already be healthier, or may exercise in other ways — both are confounders. A stronger design would compare runners to non-runners while controlling for other risk factors, or (better) randomize a running program among willing participants.
  15. (a) Randomized controlled experiment. (b) Explanatory: treatment group (categorical, three levels). Response: psychological well-being. (c) No — participants were student volunteers at one university, so results do not generalize to the broader population of young adults. (d) Yes — random assignment supports a causal claim about the studied volunteers. (e) Replace “proof” with “evidence”: the study provides evidence, not proof.

E.3 Chapter 3 — Exploring Categorical Data

Solutions to the odd-numbered exercises in Chapter 3 — Exploring Categorical Data.

  1. (a) The bar plot shows the order of the categories (by frequency) and lets us read off the relative sizes directly from the bar lengths. (b) There are no features apparent in the pie chart that are not also apparent in the bar plot. (c) The bar plot — lengths are easier for the eye to compare than angles, and the ordering makes the ranking obvious.
  2. (a) Yes, views on the protests and age appear to be associated. The stack breakpoints shift systematically across age groups: older respondents contribute more to the “strongly oppose” segment, while younger respondents lean toward the “strongly support” segment. (b) Answers may vary. Political ideology or party affiliation, education level, and primary news source could all plausibly explain part of the observed pattern.
  3. (a) The stacked (count) plot makes the number of patients in each group easy to see. (b) The standardized (proportion) plot makes the survival rate in each group easy to compare. (c) The standardized plot is the better display for this study — the scientific question is whether transplant improves survival, and comparing proportions between the two groups is exactly what that question asks.
  4. (a) Overall, about 41.0% of JetBlue flights were delayed and 40.7% of United flights were delayed — JetBlue’s overall delay rate is slightly higher. (b) By destination: SFO — JetBlue 39.7%, United 40.0%; LAX — JetBlue 40.1%, United 41.0%; BQN — JetBlue 45.7%, United 48.8%. In every city, United’s delay rate is higher. (c) This is Simpson’s paradox. A large share of JetBlue’s flights fly to BQN, where delays are high, pulling JetBlue’s overall rate up. A large share of United’s flights fly to SFO and LAX, where delays are lower, pulling United’s overall rate down. The unequal distribution of flights across cities reverses the direction of the comparison when the data are aggregated.

StatLens Exercises

  1. (a) Read the totals directly from the row and column margins of the two-way table; the transplant-and-survived cell is in the interior. A two-way table gives counts per combination, not just marginal totals. (b) Row proportions (rows = group) answer “given group, what fraction survived?” The treatment row shows a higher survival rate than the control row — yes, there is an association between receiving a transplant and survival. (c) Column proportions (columns = outcome) answer a different question: “given survival, what fraction were in the treatment group?” This is the reverse conditional — useful for a “how many of the survivors received a transplant” question, but not the cause-and-effect question we usually want to answer. Reversing the direction of conditioning changes what the number means. (d) Under independence, both group bars would show the same survived/died split — group membership would be irrelevant to outcome. The actual chart shows the treatment bar with more “survived” and less “died” than the control bar; that mismatch is the visual signature of an association.

E.4 Chapter 4 — Exploring Quantitative Data

Solutions to the odd-numbered exercises in Chapter 4 — Exploring Quantitative Data.

  1. (a) Decrease: the new score is smaller than the mean of the 24 previous scores. (b) Calculate a weighted mean, using weight 24 for the previous mean and weight 1 for the new score: \((24 \times 74 + 1 \times 64)/(24+1) = 73.6\). (c) The new score is more than one standard deviation below the previous mean, so the standard deviation increases.
  2. Any 10 employees whose average number of days off is between the minimum and the mean number of days off for the entire workforce at this plant.
  3. (a) Dist B has the higher mean since \(20 > 13\), and the higher standard deviation since 20 sits farther from the rest of the values than 13 does. (b) Dist A has the higher mean since \(-20 > -40\); Dist B has the higher standard deviation since \(-40\) is farther from the rest of its data than \(-20\) is from A’s. (c) Dist B has the higher mean since every value is larger than the corresponding value in Dist A, but both distributions have the same standard deviation since they are equally variable around their respective means. (d) Both distributions have the same mean (centered at 300), but Dist B has the higher standard deviation since its observations spread farther from the center.
  4. (a) The median is about 26. (b) Since the distribution is right-skewed, the mean is higher than the median. (c) \(Q_1\) is between 15 and 20, \(Q_3\) is between 35 and 40, so the IQR is about 20. (d) Using the \(1.5 \times \text{IQR}\) rule: upper fence \(= Q_3 + 1.5 \times \text{IQR} = 37.5 + 1.5 \times 20 = 67.5\); lower fence \(= Q_1 - 1.5 \times \text{IQR} = 17.5 - 1.5 \times 20 = -12.5\). The lowest AQI in the sample is not below 5 and the highest is not above 65, so no day in this sample would be flagged as unusually low or high.
  5. The histogram reveals that the distribution is bimodal, which the box plot hides entirely. The box plot, however, makes it easy to read precise values for observations outside the whiskers.
  6. (a) Right-skewed: there is a natural boundary at 0 and only a few people own many pets. Report the median and IQR. (b) Right-skewed: there is a natural boundary at 0 and only a few people live very far from work. Report the median and IQR. (c) Symmetric. Report the mean and standard deviation. (d) Left-skewed. Report the median and IQR. (e) Left-skewed. Report the median and IQR.
  7. No — we would expect this distribution to be right-skewed, for two reasons: hours of TV watched has a natural lower boundary at 0, and the standard deviation is very large relative to the mean, which forces a long upper tail.
  8. No. The outliers are likely to be the maximum and the minimum themselves, so a statistic based on those values (the range) cannot be robust to outliers.
  9. The 75th percentile score is 82.5, so 5 students will earn an A. By definition, exactly 25% of the class scores above the 75th percentile.
  10. (a) If \(\bar{x}/\text{median} = 1\), then \(\bar{x} = \text{median}\) — most likely for a symmetric distribution. (b) If \(\bar{x}/\text{median} < 1\), then \(\bar{x} < \text{median}\) — most likely for a left-skewed distribution, since the mean is pulled down by the low tail more than the median is. (c) If \(\bar{x}/\text{median} > 1\), then \(\bar{x} > \text{median}\) — most likely for a right-skewed distribution, since the mean is pulled up by the high tail.
  11. (a) About 68% of adults have IQ scores in \((100-15, \ 100+15) = (85, \ 115)\). (b) About 95% fall in \((100-30, \ 100+30) = (70, \ 130)\). (c) About 99.7% fall in \((100-45, \ 100+45) = (55, \ 145)\). (d) About 5% of adults fall outside \((70, 130)\), split evenly between the two tails, so about 2.5% score at or above 130.

StatLens Exercises

  1. (a) The distribution is right-skewed and unimodal: most counties have low poverty rates (roughly 5–15%), with a long tail of higher-poverty counties. (b) With only 5–6 bins, the overall shape (right-skewed) is unchanged, but subtle detail is lost — you can’t see whether the peak is smooth or bumpy, and small local clusters get merged. (c) With 30–40 bins, the shape conclusion still doesn’t change; extra detail (small local peaks and gaps) becomes visible, but at some point the extra bins add noise rather than signal. (d) The overall shape (skew, modality, center, spread) is a real feature of the data and is stable across reasonable binnings. Fine detail (exact peak location, minor bumps) depends on the bin width. As long as you don’t over- or under-bin dramatically, different histograms of the same data tell the same summary story even if they look different up close.
  2. (a) Boxplot. The median appears directly as the line inside the box — read it off. A histogram gives only an approximate median (somewhere near the middle of the mass). (b) Boxplot. Extreme values are flagged explicitly as outlier points beyond the whiskers; count how many appear above 5 million. (c) Histogram (or dotplot). Skewness is visually obvious — one long tail versus the other. A boxplot shows skew through asymmetric whiskers, but the histogram tells the story faster. (d) With 50+ bins, almost all counties collapse into the leftmost bin because county populations are extremely right-skewed — a few counties (LA, Cook) have millions, but the vast majority have fewer than 100{,}000. The histogram becomes an uninformative single tall spike. A log transformation would help; a boxplot at least shows the outliers explicitly.
  3. (a) Read the mean and standard deviation from the summary panel; NBA heights average roughly 78–79 inches with an SD of about 3–4 inches (exact values depend on the sample). (b) About 68% of players should fall within one SD of the mean, by the empirical rule. (c) Yes — for a bell-shaped distribution, roughly two-thirds of the data lie inside \((\bar{x} - s, \ \bar{x} + s)\). Not exact, but close. (d) About 95% within two SDs. Looking at the histogram, nearly all NBA player heights fall inside \((\bar{x} - 2s, \ \bar{x} + 2s)\). (e) No. The empirical rule requires an approximately bell-shaped distribution; Ames home prices are strongly right-skewed, so the 68%/95%/99.7% percentages don’t apply. Reporting “68% within one SD” would substantially misrepresent the data.

E.5 Chapter 5 — Relationships Between Variables

Solutions to the odd-numbered exercises in Chapter 5 — Relationships Between Variables.

  1. (a) The ridge plots show meat consumption and life expectancy separately by income group, not directly against each other. From these plots we can see that the high-income group has both the highest meat consumption and the highest life expectancy, but we cannot see the meat vs. life-expectancy relationship within an income group, so we cannot claim they are associated. (b) When a relationship is confounded, we cannot pin down the causal mechanism. A longer life expectancy in high-meat countries could be driven by higher income (which brings better medical care, sanitation, nutrition, and many other life-extending factors) rather than by meat itself. (c) Break the data into subgroups by the confounding variable (income), then examine the relationship of interest (meat vs. life expectancy) separately within each subgroup.
  2. (a) Positive association: mammals with longer gestation periods tend to live longer. (b) The association would still be positive; swapping the axes does not change the direction of a relationship. (c) No, they are not independent — as (a) describes, they trend together.
  3. The plot should show a ramp-up: an early period of nearly exponential growth (few cells reproducing freely), followed by leveling off as the population approaches the one-million-cell carrying capacity. The resulting curve is S-shaped (logistic), rising quickly at first and then flattening out.
  4. Both distributions are unimodal and right-skewed, with a long tail of older winners. The center of the best-actor distribution sits noticeably above the best-actress distribution — best-actor winners are, on average, roughly 7 years older than best-actress winners. Spread is comparable, though the best-actress distribution shows slightly more variability. In short: same shape, similar spread, but best-actor winners tend to be older.
  5. (a) The manual transmission group typically shows the higher median city MPG, while the automatic group has the larger IQR — the automatic group spans a wider variety of vehicle types (from small commuters to large SUVs), so it is more variable. (b) Both distributions are unimodal and right-skewed: most vehicles cluster in the moderate-MPG range, with a long tail of very fuel-efficient outliers (small hybrids and specialty vehicles). (c) The density curves overlap substantially in the middle range. The difference between the two group medians is small relative to the within-group variability, so transmission type alone does not explain most of the differences in city MPG. (d) A complete comparison names all four features: shape (both right-skewed, unimodal), center (manual has slightly higher median), spread (automatic has larger IQR), and unusual features (a few high-MPG outliers in both groups, likely hybrids or small cars). Overall: a modest difference in center against a backdrop of substantial overlap.

E.6 Chapter 6 — Sampling Variability

Solutions to the odd-numbered exercises in Chapter 6 — Sampling Variability.

  1. (a) Mean — each student reports a numerical value (hours). (b) Mean — each student reports a number (a percentage), and averaging is meaningful. (c) Proportion — each student answers Yes/No, a categorical response. (d) Mean — each student reports a percentage, like in (b). (e) Proportion — each recent graduate reports whether they expect a job within a year, a categorical response.
  2. (a) The center stays at \(\mu\), the population mean. The sample mean is unbiased, so its long-run average equals \(\mu\) regardless of \(n\). Bigger samples don’t shift the target — they hit it more precisely. (b) The spread shrinks. SE\((\bar{x}) = \sigma/\sqrt{n}\), so quadrupling \(n\) from 25 to 100 cuts the SE roughly in half; going from 25 to 400 cuts it to about a quarter. (c) The shape becomes more symmetric and bell-shaped. At \(n = 25\) the sampling distribution can still look lumpy or skewed (depending on the population); by \(n = 400\) it typically looks like a smooth mound. (d) \(n = 400\) gives the most precise estimate but requires the most data collection effort. \(n = 25\) is cheap but has \(4\times\) the SE of \(n = 400\). The typical study picks the smallest \(n\) that meets the precision target the research question demands.
  3. (a) Data distribution — each observation is one patient’s age; no averaging across samples. (b) Sampling distribution — each observation is the mean of a sample of 30, a statistic. (c) Sampling distribution — each observation is one person’s sample proportion \(\hat{p}\); the collection of 400 \(\hat{p}\) values is the sampling distribution. (d) Data distribution — each observation is one student’s SAT score. (e) The sampling distribution is narrower. Averaging (or otherwise summarizing) many observations dampens out extremes; SE\((\bar{x}) = \sigma/\sqrt{n}\), which is smaller than \(\sigma\) once \(n > 1\).
  4. (a) \(s = 1.4\) hours. “A typical student” is about individual variation — the sample standard deviation. (b) SE\((\bar{x}) = 0.07\) hours. “How much would the sample mean vary from sample to sample” is exactly the standard error. (c) \(s\) — variability across students is the standard deviation of the data. (d) Averaging 400 observations damps out individual variability: SE\((\bar{x}) = s/\sqrt{n} = 1.4/\sqrt{400} = 0.07\). Dividing by \(\sqrt{n} = 20\) shrinks the spread by a factor of 20 relative to \(s\).
  5. Match: A \(= (n=100, p=0.20)\) — centered near 0.20, narrow. B \(= (n=25, p=0.50)\) — centered near 0.50, wide. C \(= (n=25, p=0.20)\) — centered near 0.20, wide. D \(= (n=100, p=0.50)\) — centered near 0.50, narrow. The center told you \(p\) (which histograms cluster near 0.20 vs. 0.50); the spread told you \(n\) (wider = smaller \(n\); quadrupling \(n\) from 25 to 100 roughly halves the spread).

StatLens Exercises

  1. (a) Population = all 14,000 undergraduates; sample = the 200 surveyed. (b) A statistic, \(\hat{p} = 0.66\) — computed from the sample. (c) A parameter, \(p\) — the fraction of all 14,000 undergraduates who work 10+ hours. Its value is unknown; we would have to survey all 14,000 to know it. The statistic \(\hat{p}\) is our best guess at \(p\).
  2. (a) At \(n = 5\), the sampling distribution of \(\bar{x}\) is less skewed than the population but not yet symmetric — still noticeably right-skewed and fairly wide. (Many students predict it will look just like the population; it doesn’t.) (b) Larger \(n\) changed the spread. Drawing more samples only fills in the same sampling distribution more completely — its spread does not shrink. Increasing \(n\) produces a different, narrower sampling distribution. These two knobs are easy to confuse. (c) The standard error is the standard deviation of the sampling distribution. The observed SD of many simulated sample means estimates that same quantity, so with enough draws the two should nearly agree.
  3. (a) The student confused the standard error with the standard deviation of the data. An SE of 0.3 describes how much the sample mean would vary from sample to sample — not how individual students vary. Individual sleep times vary far more than \(\pm 0.6\) hours. (b) The interval \(7.1 \pm 2(0.3) = (6.5, 7.7)\) is an approximate 95% confidence interval for the mean sleep time — a range of plausible values for \(\mu\), not a range that captures 95% of individual students.

E.7 Chapter 7 — Randomization Tests

Solutions to the odd-numbered exercises in Chapter 7 — Randomization Tests.

  1. (a) Alternative — claims a directional predictive relationship. (b) Null — “on average run the same speed” asserts no difference. (c) Alternative — claims the probability is larger than 0.2. (d) Alternative — claims the mean has changed. (e) Null — “risk is equal” asserts no difference. (f) Alternative — claims caffeine affects mean birth weight. (g) Null — “the probability is the same” asserts no difference.
  2. (a) In words: \(H_0\): On average, New Yorkers sleep 8 hours per night. \(H_A\): On average, New Yorkers sleep less than 8 hours per night. In symbols: \(H_0: \mu = 8\) vs. \(H_A: \mu < 8\), where \(\mu\) is the mean nightly sleep of New Yorkers. (b) In words: \(H_0\): Employees spend, on average, 15 minutes per business day on non-work activities. \(H_A\): Employees spend more than 15 minutes, on average, on non-work activities during March Madness. In symbols: \(H_0: \mu = 15\) vs. \(H_A: \mu > 15\), where \(\mu\) is the mean non-work minutes per day during March Madness.
  3. (a) (i) False. The two groups are very different sizes (67,593 vs. 159,978), so raw counts are not comparable. Compare the proportions of cardiovascular events in each group. (ii) True. The rate is \(2{,}593/67{,}593 \approx 3.8\%\) for Rosiglitazone and \(5{,}386/159{,}978 \approx 3.4\%\) for Pioglitazone. (iii) False. This is an observational study, so we cannot conclude causation from a difference in rates alone. (iv) True. With only the two-way table we cannot separate a genuine treatment effect from ordinary sampling variability. (b) Overall rate of cardiovascular problems: \(7{,}979/227{,}571 \approx 0.035\). (c) Under independence we would expect \(67{,}593 \times (7{,}979/227{,}571) \approx 2{,}370\) patients in the Rosiglitazone group to experience cardiovascular problems. (d) (i) \(H_0\): Treatment and cardiovascular problems are independent — the observed 3.8% vs. 3.4% gap is due to chance. \(H_A\): The two are not independent — Rosiglitazone is associated with higher risk of cardiovascular problems. (ii) A higher proportion in the Rosiglitazone group would give more support for \(H_A\). (iii) The randomization histogram is centered near 0 and its 100 simulated differences never reach the observed difference of about \(0.038 - 0.034 = 0.004\); almost none of the shuffles produce a gap this large. That gives strong evidence against independence and supports the alternative — the association between Rosiglitazone and cardiovascular problems is very unlikely to be a chance finding.

StatLens Exercises

  1. (a) The parameter is \(\mu_T - \mu_C\), the difference in population mean exam scores between the tutoring and control conditions. The observed 4.2-point difference in sample means is our estimate. (b) \(H_0: \mu_T - \mu_C = 0\) vs. \(H_A: \mu_T - \mu_C > 0\). The test is one-sided because the question asks whether tutoring improves scores — a directional claim stated in advance. (c) Under \(H_0\) every student’s score would be the same regardless of group label, so the labels are interchangeable. Shuffling the labels generates datasets consistent with that “no-effect” world. Random assignment is what makes the shuffle valid: students were placed into groups by chance, not self-selection, so under \(H_0\) the labels really are arbitrary.
  2. (a) The null distribution is centered at 0 — the value \(H_0\) asserts. It must be, because the shuffling procedure breaks any real association between group label and outcome; averaged over many shuffles the differences cancel. (b) The p-value equals (number of shuffled differences at least as extreme as the observed one) \(/\,1000\). StatLens reports this directly; for the sex-discrimination data it lands around 0.02–0.03 one-sided. A small p-value means the observed difference is rare in the “no-effect” world, so the data give evidence against \(H_0\). (c) If the observed difference sat near the middle of the null distribution, roughly half of the shuffles would be at least that extreme, so the p-value would be near 0.5. We would fail to reject \(H_0\) — the observed gap is exactly the kind of fluctuation chance alone routinely produces.
  3. (a) With enough shuffles the null distribution is roughly bell-shaped, centered at 0 by construction, and its spread equals the standard error of \(\bar{x}_1 - \bar{x}_2\). (b) The standard deviation of the null distribution is the simulated standard error of \(\bar{x}_1 - \bar{x}_2\). The formula \(\text{SE}(\bar{x}_1 - \bar{x}_2) = \sqrt{s_1^2/n_1 + s_2^2/n_2}\) is an analytical estimate of the same quantity, derived under independence and large-sample assumptions. When both are valid, they agree — which is why a randomization test and a \(t\)-test on large, well-behaved samples give nearly identical p-values. (c) With small \(n\) and an outlier the null distribution is discrete (only finitely many label permutations exist), often visibly non-normal in the tails, and may show spikes or gaps. The smooth \(t\) distribution can over- or under-state the tail probability depending on how the outlier moves around under permutations, so the two p-values can differ substantially. (d) The claim is backwards. The randomization test needs no normality assumption — it just enumerates the shuffling distribution that actually applies under random assignment. The \(t\)-test relies on the CLT to make \(\bar{x}_1 - \bar{x}_2\) approximately normal, and that approximation is weakest when \(n\) is small. So the randomization test is the safer choice in small samples, not the reverse.

E.8 Chapter 8 — Decision Errors

Solutions to the odd-numbered exercises in Chapter 8 — Decision Errors.

  1. (a) \(H_0\): anti-depressants do not affect the symptoms of Fibromyalgia. \(H_A\): anti-depressants do affect the symptoms of Fibromyalgia (either helping or harming). (b) Concluding that anti-depressants either help or worsen Fibromyalgia symptoms when they actually do neither. (c) Concluding that anti-depressants do not affect Fibromyalgia symptoms when they actually do.
  2. (a) True. A 99% CI is wider than a 95% CI computed from the same sample (larger \(t^*\) or \(z^*\)), so it contains every value the 95% interval contains. (b) False. Decreasing \(\alpha\) decreases the Type I error rate — that is what \(\alpha\) is. Lowering \(\alpha\) makes rejection harder and so raises the Type II error rate \(\beta\) at a fixed sample size. (c) False. Failing to reject \(H_0\) does not prove \(H_0\) is true; it only means the evidence against \(p = 0.5\) was not strong enough. The true \(p\) could be 0.5, or close to 0.5, or the test could simply have lacked power. (d) True. With very large \(n\), the standard error becomes tiny, so even a small gap between the null value and the point estimate produces a large test statistic and a small \(p\)-value. Statistical discernibility does not imply the effect is practically important.
  3. Seven confidence intervals out of 100 missing \(\pi\) is completely consistent with either a 95% or a 90% confidence level. At 95% we would expect about 5 misses in 100 — and 7 is well within the sampling noise you would see even when the students did everything correctly. (At 90% we would expect about 10 misses, so 7 is again unremarkable.) The professor should not dock the seven students for missing \(\pi\): an occasional CI that misses the parameter is exactly what a confidence interval is defined to allow.
  4. In every case, \(p\) is the population proportion described in the scenario, and \(H_0\) takes an equal sign.

    \(a\) $H_0: p = 0.20 \quad H_A: p > 0.20$. Here $p$ is the proportion of the 30 pre-test questions a student would answer correctly *in the long run*; random guessing on 5-option questions gives $p = 1/5 = 0.20$, and the professor wants to detect whether students perform better than chance.
    (b) $H_0: p = 0.32 \quad H_A: p \neq 0.32$. Here $p$ is the proportion of patients on the new intervention who would experience reduced blood pressure. The trial is looking for a *different* result, not necessarily better, so the alternative is two-sided.
    (c) $H_0: p = 0.67 \quad H_A: p > 0.67$. Here $p$ is the proportion of registered voters who will turn out for the next presidential election. The question asks specifically about *higher* turnout, so the alternative is one-sided.</li>

:::

StatLens Exercises

  1. (a) \(H_0: p = 0.50\) versus \(H_A: p > 0.50\). (b) Type I error: we conclude the drug raises the success rate when it really doesn’t. Type II error: we conclude the drug doesn’t help when it actually does. (c) Patients bear the cost of a Type I error — they receive an ineffective drug (possibly with side effects) in place of a known treatment. Patients and the company bear the cost of a Type II error — a real cure goes undeveloped. When a Type I error has serious safety consequences, regulators demand a small \(\alpha\) (often 0.025 or 0.01 one-sided), which raises the evidence bar for rejecting \(H_0\).
  2. (a) About 5% of the 500 studies reject — by construction, that is what \(\alpha = 0.05\) means. Every rejection here is a Type I error because \(H_0\) is true. (b) About 1%. The Type I error rate is whatever you set \(\alpha\) to (given the test’s conditions hold); lowering \(\alpha\) from 5% to 1% cuts the false-alarm rate fivefold. (c) About 88–93% reject — this is the power of the test at \(p_{\text{true}} = 0.65\), \(n = 100\), \(\alpha = 0.05\). The non-rejections (roughly 7–12% of studies) are Type II errors: \(H_A\) is true and we failed to detect it. (d) Power rises with \(n\): more data makes it easier to distinguish \(p = 0.65\) from the null value 0.50. Correspondingly, \(\beta\) falls as \(n\) grows.
  3. (a) & (b) At \(n = 100\), \(\Delta p = 0.15\), and \(\alpha = 0.05\) the simulated power is roughly 88–93% — comfortably above the 80% bar, so this trial is adequately powered. A 15-percentage-point effect is large enough that 100 patients detect it most of the time; students often underestimate the power available at large effects. (A formal sample-size calculation, coming in Ch 31, gives the precise number.) (c) About a 10% chance the trial fails to reject \(H_0\) even though the drug really works — a Type II error. A non-significant result is never proof the drug does nothing; even well-powered trials carry a residual miss rate. (d) Two levers: (i) increase \(n\) — larger samples raise power for any fixed effect size and \(\alpha\); this is usually under the researcher’s control (budget permitting). (ii) raise \(\alpha\) — accepting more false positives raises the true-positive rate too; this is usually not under the researcher’s control once the field has a convention. Other levers include increasing the effect size (enroll patients with more room to improve), decreasing variability (tighter inclusion criteria), or switching to a one-sided test when the direction is truly known in advance.

E.9 Chapter 9 — Bootstrap Confidence Intervals

Solutions to the odd-numbered exercises in Chapter 9 — Bootstrap Confidence Intervals.

  1. (a) The statistic is the observed sample proportion of outdoor videos, \(37/128 \approx 0.289\). The parameter is the true proportion of all YouTube videos that take place outdoors — unknown, since we haven’t surveyed the entire population. (b) The statistic is \(\hat{p}\); the parameter is \(p\). (c) The proportion of “outdoors” videos in each bootstrap resample — a bootstrap sample proportion. (d) Near the observed \(\hat{p} \approx 0.289\); a bootstrap distribution is centered at the sample statistic, not at the population parameter. (e) Reading the middle 90% of the histogram gives roughly \((0.22,\, 0.35)\). (f) We are 90% confident that between about 22% and 35% of all YouTube videos take place outdoors.
  2. With 98% confidence, the true proportion of U.S. adults (in 2022) who get news from social media at least sometimes is between about \(0.487\) and \(0.510\). Note that the width is only about 2 percentage points — a consequence of the very large sample size (\(n = 12{,}147\)).
  3. Match each histogram to the sample proportion it was built from by finding the histogram whose center is nearest that \(\hat{p}\) value (a bootstrap distribution is centered at the observed sample statistic). (a) \(\hat{p} = 0.13 \to\) A; (b) \(\hat{p} = 0.22 \to\) B or D (both are near 0.2–0.25); (c) \(\hat{p} = 0.30 \to\) C; (d) \(\hat{p} = 0.43 \to\) has no strong match among these four (the histograms center below 0.35); (e) \(\hat{p} = 0.75 \to\) none.
  4. (a) Supported. The entire interval (54%, 64%) lies above 50%, so the data are consistent with a majority. (b) Not supported. The value 70% lies above the upper endpoint of 64%, so the interval provides evidence against the researcher’s conjecture. (c) A 90% CI is narrower than the 95% CI — both are centered in the same place, and 70% already sits outside the wider 95% interval, so it will also fall outside the narrower 90% interval. The conjecture would still be rejected.
  5. (a) Both methods agree. The SE method assumes an approximately normal bootstrap distribution (statistic \(\pm 2\,\text{SE}\)). When the bootstrap distribution is symmetric and bell-shaped, that assumption holds and the two methods give nearly identical intervals. (b) Percentile is more trustworthy. The SE method centers the CI symmetrically around \(\bar{x}\), but a right-skewed distribution has an asymmetric middle 95% — the upper tail is longer. Percentile captures that asymmetry; SE forces a symmetry that isn’t in the data. (c) Percentile is more trustworthy. The SE method could push the upper endpoint above 1 (impossible for a proportion). Percentile respects the natural boundary because it just reports what’s in the middle 95% of the bootstrap draws. (d) Neither is fully trustworthy — a lumpy bootstrap distribution suggests \(n\) is too small for the bootstrap to work well. If you must report an interval, percentile is safer than SE because it doesn’t require the bootstrap distribution to be roughly normal. (e) The SE method assumes the bootstrap distribution is approximately symmetric and bell-shaped, so that statistic \(\pm 2\,\text{SE}\) captures the middle 95%. Percentile makes no shape assumption — it just reads off the 2.5th and 97.5th percentiles. When the bootstrap distribution is skewed, bounded, or lumpy, percentile is the safer default.

StatLens Exercises

  1. (a) The parameter is the population median wait time at this drive-through, with symbol \(M\) (sometimes \(\tilde{\mu}\)). The sample median \(\tilde{x}\) of the 80 recorded times is the estimate — not the parameter. (b) The mean has a tidy formula for its standard error (\(s/\sqrt{n}\)) because the CLT gives \(\bar{x}\) an approximately normal sampling distribution. The median has no clean closed-form SE, and its sampling distribution can be skewed and irregular. The bootstrap needs no formula — it resamples and shows the sampling distribution, so the median works as easily as the mean. (c) Use the percentile method when the bootstrap distribution is skewed (a symmetric \(\hat{\theta} \pm z^*\,\text{SE}\) would put the interval in the wrong shape). Use the SE method when the bootstrap distribution is roughly symmetric and bell-shaped, where the two methods coincide and the SE form is conceptually closer to the formulas that appear later. When in doubt, default to percentile.
  2. (a) Right-skewed, centered near \(\hat{p} \approx 0.048\), with narrow spread (small SE). The shape is right-skewed because \(\hat{p}\) is small: resamples cannot go below 0 (a hard floor) but can drift somewhat higher, so the tail extends to the right. (b) The 95% percentile interval will land near \((0.01,\, 0.11)\) — typically just barely covering the national rate of 0.10. Because the interval contains 0.10, the data do not provide strong evidence that this consultant’s true complication rate differs from the national average. (c) A symmetric interval \(\hat{p} \pm 2\,\text{SE}_{\text{boot}}\) is centered at \(\hat{p}\) and the same width on both sides, which is the wrong shape for a right-skewed distribution. Its lower endpoint could even go negative (below 0!), which is impossible for a proportion. The percentile interval respects the skew naturally; the symmetric formula imposes a shape the data don’t have.
  3. (a) For well-behaved, moderate-\(n\) data the two 95% intervals are typically very close — agreeing to within a small fraction of the SE. The \(t\)-procedure’s assumptions (independent observations, approximate normality of \(\bar{x}\)) are met, so the formula and the bootstrap converge. (b) They differ noticeably on small skewed data. The \(t\)-interval is forced to be symmetric about \(\bar{x}\) (equal width on each side, controlled by \(t^*\cdot s/\sqrt{n}\)); the bootstrap percentile interval can be asymmetric, following the actual skew of the bootstrap distribution. The midpoint can shift and the width on one side can exceed the other. (c) The \(t\)-interval is built from a symmetric distribution (the \(t\)), so it can only express symmetric uncertainty around \(\bar{x}\). When the sampling distribution of \(\bar{x}\) is itself skewed — which happens with small \(n\) and a skewed population — the symmetric interval is the wrong shape, even if its center is right. The bootstrap imposes no shape; it shows what the sampling distribution actually looks like. In well-behaved data the two shapes coincide, so the intervals agree. (d) The intervals aren’t “disagreeing about a fact” — they are two approximations to the same target, and they converge as the conditions improve. When they diverge, it is a signal that the \(t\)-interval’s symmetry assumption has failed, not that one of them computed wrong. Trust the bootstrap in that case: it makes no normality or symmetry assumption, so it stays honest exactly where the formula breaks down. This mirrors the logic from the proportions chapter: simulation is the more general tool; the formula is the shortcut that only applies when its conditions are met.

E.10 Chapter 10 — Sampling Distributions via Simulation

Solutions to the odd-numbered exercises in Chapter 10 — Sampling Distributions via Simulation.

  1. (a) Sampling distribution of the sample proportion \(\hat{p}\). (b) Approximately symmetric. We’re told the true proportion is somewhere in \([0.05,\,0.30]\), and with \(n = 800\) the success–failure condition (\(np \ge 10\) and \(n(1-p) \ge 10\)) holds comfortably across that range, so the sampling distribution is bell-shaped. (c) Standard error of \(\hat{p}\). (d) With only \(n = 250\) homes per sample, the sampling distribution will be more variable (wider) than with \(n = 800\). Standard error scales as \(1/\sqrt{n}\), so shrinking \(n\) from 800 to 250 inflates the SE by a factor of \(\sqrt{800/250} \approx 1.79\).
  2. (a) (i) and (iii) are true. The CLT is a statement about the sampling distribution of \(\bar{X}\); it does not change the population distribution or the population mean \(\mu\). (b) Mean \(= \mu = 3.15\). Standard error \(= \sigma/\sqrt{n} = 0.60/\sqrt{400} = 0.030\). (c) The CLT is about the sampling distribution of \(\bar{X}\) — the distribution of many sample means, one per sample — not about the data in any single sample. A single sample from a skewed population is still skewed no matter how large \(n\) is; only the distribution of means across samples becomes bell-shaped. (d) A common rule of thumb is \(n \ge 30\) for populations that are not too skewed. Strongly skewed or heavy-tailed populations may need \(n = 100\) or more. There is no single magic number: the rougher the population, the larger \(n\) needs to be.

StatLens Exercises

  1. (a) Most students predict (ii). (b) In practice you’ll see (iii): at \(n = 5\) from a strongly right-skewed population, the sampling distribution of \(\bar{x}\) still carries visible skew — the CLT hasn’t kicked in yet. (c) By \(n = 30\) the skew is usually mostly gone for a moderately skewed population; by \(n = 100\) the distribution looks essentially normal. (d) A bimodal population typically needs a larger \(n\) than a right-skewed one, because the two modes have to “average out” before the bell shape emerges. The CLT still eventually wins, but “\(n = 30\)” is a rule of thumb, not a guarantee — the shape of the population dictates how fast the sampling distribution becomes normal.
  2. (a) Predicted center \(= p = 0.20\). Predicted spread \(= \sqrt{p(1-p)/n} = \sqrt{(0.20)(0.80)/20} = \sqrt{0.008} \approx 0.089\). (b) With 5000 samples the lab should report a center very close to 0.20 and an SD of \(\hat{p}\) very close to 0.089. (c) Right-skewed. With \(p = 0.20\) and \(n = 20\), \(\hat{p}\) has a hard floor at 0 and a long tail toward larger values; the success–failure check gives \(np = 4 < 10\), so the normal approximation is poor here. (d) At \(n = 200\), \(np = 40\) and \(n(1-p) = 160\) — both well above 10 — so the sampling distribution is now approximately normal and much tighter, with SE \(= \sqrt{(0.20)(0.80)/200} \approx 0.028\) (about one-third the earlier spread).
  3. (a) Students at one table are not independent of each other: they’re likely clustered in friend groups, dietary preferences, or ordering patterns. The CLT’s SE formula assumes i.i.d. observations, so it underestimates the actual variability of \(\bar{x}\); a CLT-based CI will under-cover the true mean (real coverage far below 95%). (b) The Cauchy distribution is the classic CLT counterexample: it has no defined mean or variance, and the sampling distribution of \(\bar{x}\) stays Cauchy at every \(n\) — it never becomes normal, no matter how large \(n\) gets. (c) Truly infinite-variance populations are rare in practice; subtle clustering (same dorm, same city, same day) is everywhere. Independence is the assumption to worry about, not the existence of finite variance. (d) Not very comforting. When the distribution is heavy-tailed enough that a single viral post can drive \(\bar{x}\), the CLT still applies “in the limit” but the limit is approached very slowly — the practical \(n\) needed for normality may be in the millions, not the thousands.

E.11 Chapter 11 — Normal Approximation

Solutions to the odd-numbered exercises in Chapter 11 — Normal Approximation.

  1. (a) The general formula is \(\text{point estimate} \pm z^\star \times SE\), with point estimate \(= 45\%\), \(z^\star = 1.96\) for 95% confidence, and \(SE = 1.2\%\): \(45\% \pm 1.96 \times 1.2\% \to (42.6\%, 47.4\%)\). We are 95% confident that the proportion of U.S. adults who live with one or more chronic conditions is between 42.6% and 47.4%. (b) (i) False. A 95% confidence interval will “miss” the true value about 5% of the time. (ii) True. Long-run coverage: about 950 of 1,000 such intervals would capture the true proportion. (iii) True. 50% lies outside the interval, so the null value \(p = 0.5\) would be rejected at \(\alpha = 0.05\). (iv) False. Standard error describes the sampling variability of the estimate, not individuals’ uncertainty about their own answers.
  2. (c) is the correct interpretation: a Z score of 0.47 means the sample proportion is 0.47 standard errors greater than the hypothesized value of the population proportion. The others describe different ideas: (a) confuses the \(Z\) score with a probability that \(H_0\) is true; (b) describes a p-value, not a \(Z\) score; (d) confuses the sample proportion itself with a scaled distance; (e) drops the units — it’s “0.47 SEs away,” not “0.47 units away”; (f) claims \(\hat{p} = 0.47\), which is unrelated.
  3. (a) \(P(Z < -1.35) \approx 0.0885\), or 8.85%. (b) \(P(Z > 1.48) \approx 0.0694\), or 6.94%. (c) \(P(-0.4 < Z < 1.5) = P(Z<1.5) - P(Z<-0.4) \approx 0.9332 - 0.3446 = 0.5886\), or about 58.9%. (d) \(P(|Z| > 2) = P(Z<-2) + P(Z>2) \approx 2(0.0228) = 0.0455\), or about 4.6%.
  4. (a) Verbal Reasoning: \(N(\mu = 151, \sigma = 7)\); Quantitative Reasoning: \(N(\mu = 153, \sigma = 7.67)\). (b) \(Z_{VR} = (160 - 151)/7 = 1.29\); \(Z_{QR} = (157 - 153)/7.67 = 0.52\). (c) Sophia scored 1.29 standard deviations above the mean on Verbal Reasoning and 0.52 standard deviations above the mean on Quantitative Reasoning. (d) She did better on Verbal Reasoning — her \(Z\) score is higher there. (e) \(Perc_{VR} \approx 0.90\) (90th percentile); \(Perc_{QR} \approx 0.70\) (70th percentile). (f) About 10% of test takers did better than Sophia on Verbal Reasoning and about 30% did better on Quantitative Reasoning. (g) Raw scores are on different scales (different means and SDs), so a raw comparison is misleading; percentiles put both sections on a common footing. (h) The \(Z\) scores in part (b) would not change (a \(Z\) score is just a rescaling), but parts (d)–(f) would no longer be answerable — percentiles require a distributional model.
  5. (a) The 80th percentile of \(N(153, 7.67)\): \(z = 0.84\), so the score is \(153 + 0.84 \times 7.67 \approx 159\). (b) “Worse than 70% of test takers” is the 30th percentile of \(N(151, 7)\): \(z = -0.52\), so the score is \(151 + (-0.52) \times 7 \approx 147\).
  6. (a) \(Z = (83 - 77)/5 = 1.2\), and \(P(Z > 1.2) \approx 0.1151\) — about an 11.5% chance of an 83°F day or hotter. (b) The 10th percentile: \(z = -1.28\), so the cutoff is \(77 + (-1.28)(5) \approx 70.6\)°F. The coldest 10% of June days in LA have highs of about 70.6°F or lower.
  7. (a) Converting via \(C = (F - 32) \cdot 5/9\): the mean becomes \((77 - 32) \cdot 5/9 = 25\)°C and the SD scales the same way, \(5 \cdot 5/9 \approx 2.78\)°C. So \(N(25, 2.78)\). (b) \(Z = (28 - 25)/2.78 \approx 1.08\); \(P(Z > 1.08) \approx 0.1401\). (c) The answer is very close to part (a) of the previous exercise (about 11.5%). Not surprising: only the units changed. The tiny difference is because 28°C corresponds to 82.4°F, not exactly 83°F. (d) Using \(Q_1 = 25 + (-0.674)(2.78) \approx 23.13\)°C and \(Q_3 = 25 + (0.674)(2.78) \approx 26.87\)°C, the IQR is \(\approx 3.73\)°C.
  8. (a) Panel C (\(n = 200\)). Larger \(n\) makes the sampling distribution smoother, more symmetric, and more sharply peaked — the normal curve hugs it closely. (b) The \(n = 10\) histogram is visibly right-skewed and bumps into the lower boundary at 0. The success–failure condition fails: \(np = 10(0.30) = 3 < 10\). (c) Panel A: \(np = 3\), \(n(1-p) = 7\)fails. Panel B: \(np = 15\), \(n(1-p) = 35\)passes. Panel C: \(np = 60\), \(n(1-p) = 140\)passes clearly. (d) The normal approximation is useful because sampling distributions (not raw data, not populations) tend toward normal as \(n\) grows. It requires enough sample size for skewness and discreteness to smooth out; when the success–failure condition fails, prefer a simulation approach.

StatLens Exercises

  1. (a) Area to the right of \(z = 1\) is about 0.159; the area to the left of \(z = -1\) is also 0.159. They’re equal because the standard normal is symmetric about 0. (b) About 0.954 for \(-2 \le z \le 2\) — consistent with “95%” within \(\pm 2\) SDs (the empirical rule’s \(\pm 2\) figure is really \(\pm 1.96\); both round to 95%). (c) \(z \approx -1.96\). This is the critical value for a 95% confidence interval; you’ll use \(z^\star = 1.96\) throughout the CI and two-sided \(z\)-test material at \(\alpha = 0.05\). (d) Total area under any density curve is 1 by definition (the value is somewhere). Area = 0 corresponds to an event like “\(z\) equals exactly 0.5” — continuous distributions assign probability 0 to any single point; only intervals carry positive probability.
  2. (a) Mean \(= 520\), standard error \(= 110/\sqrt{25} = 22\), shape approximately normal. (b) \(z = (550 - 520)/22 \approx 1.36\), so \(P(\bar{x} \ge 550) \approx 0.087\). (c) For an individual student, \(z = (550 - 520)/110 \approx 0.27\), so \(P(X \ge 550) \approx 0.39\)much bigger. Individual scores vary a lot; sample means vary less because averaging dampens individual variability, making the same 30-point gap “more surprising” for a mean than for one student. (d) Smaller. Doubling \(n\) shrinks the SE to \(110/\sqrt{50} \approx 15.6\). The same 30-point gap becomes \(z \approx 1.93\) (instead of 1.36) — farther in the tail, lower probability.
  3. (a) For well-behaved data (\(\hat{p}\) near 0.5, \(n\) large) the two 95% intervals are very close — typically agreeing to within a fraction of a percentage point. (b) They differ noticeably. The \(z\)-interval is symmetric about \(\hat{p}\); the bootstrap interval is asymmetric, following the shape of the bootstrap distribution. (c) At small \(\hat{p}\) the bootstrap distribution is right-skewed\(\hat{p}\) has a hard floor at 0, but resamples can drift upward. The \(z\)-interval forces symmetry that doesn’t match. (d) The two intervals are not both “right.” The \(z\)-interval under-covers in the troublesome case because its symmetry assumption fails. The bootstrap is more honest: its long-run coverage matches the advertised 95% because it reflects the actual shape of the uncertainty. “Honest” here means the interval’s real coverage matches its stated confidence level.

E.12 Chapter 12 — Inference for Proportions

Solutions to the odd-numbered exercises in Chapter 12 — Inference for Proportions.

  1. First, the hypotheses should be about the population proportion (\(p\)), not the sample proportion \(\hat{p}\). Second, the null value should be the claimed value (0.25), not the observed value (0.29). Correct setup: \(H_0: p = 0.25\) and \(H_A: p > 0.25\).
  2. (a) \(H_0: p = 0.20\), \(H_A: p > 0.20\). (b) \(\hat{p} = 159/650 = 0.245\). (c) Represent each respondent with a card: use 100 cards, 20 black (support “defund”) and 80 red (do not). Shuffle and draw 650 cards with replacement, calculate \(\hat{p}_{\text{sim}}\), and repeat many times (using software in practice). The simulated null distribution shows what \(\hat{p}\) values are typical if \(p = 0.20\). The p-value is the proportion of simulations with \(\hat{p}_{\text{sim}} \geq 0.245\). (d) The p-value is approximately 0.001 (only about 1 in 1,000 simulated proportions reaches 0.245 or higher). Because \(p\text{-value} < 0.05\), we reject \(H_0\): the data provide convincing evidence that more than 20% of Seattle adults support proposals to defund police departments.
  3. (a) \(H_0: p = 0.5\), \(H_A: p \ne 0.5\). (b) The p-value is roughly 0.4. There is not enough evidence to conclude that cats prefer one shape over the other — but note that with only 7 cats there is very little power to detect a preference even if one exists.
  4. (a) \(SE(\hat{p}) \approx 0.189\). (b) Roughly 0.188 (matches the theoretical SE closely). (c) Yes — the simulation-based null distribution and the normal model agree. (d) No — with only \(n = 7\) cats, the simulated \(\hat{p}\) values are visibly discrete (only a handful of distinct outcomes are possible). (e) The bootstrap draws are discrete (only a few distinct proportions occur for such a small sample), while the mathematical normal model is continuous.
  5. (a) The null hypothesis distribution assumes \(p = 0.7\); the bootstrap distribution is built from the data, which had \(\hat{p} = 0.6\). (b) The null distribution is centered at 0.7; the bootstrap distribution is centered at 0.6. (c) Both distributions have a standard error of roughly 0.1. (d) Both are reasonably symmetric. The null distribution is slightly more skewed (left) than the bootstrap because 0.7 is closer to the boundary at 1, and the boundary compresses the right tail.
  6. (a) The null-hypothesis simulation is used for hypothesis testing; the data-bootstrap distribution is used for confidence intervals. (b) \(H_0: p = 0.7\) vs. \(H_A: p \ne 0.7\). The p-value \(> 0.05\), so there is no evidence that the proportion of full-time statistics majors who work at least 5 hours per week differs from 70%. (c) We are 98% confident that the true proportion of full-time statistics majors who work at least 5 hours per week is between 35% and 80%. (d) Using \(z^\star = 2.33\), the 98% CI is approximately \((0.367, 0.833)\).
  7. (a) False. The success-failure condition fails (only \(8\%\) of a small sample), so the sampling distribution will not be approximately normal. (b) True. Because \(p = 0.08\) is close to 0 and \(\hat{p}\) cannot go below 0, the sampling distribution is right-skewed for small samples. (c) False. \(SE_{\hat{p}} = 0.0243\), and \(\hat{p} = 0.12\) is \((0.12 - 0.08)/0.0243 \approx 1.65\) SEs away — not unusual. (d) True. \(\hat{p} = 0.12\) is \(2.32\) SEs from \(p\), which is unusual. (e) False. Doubling \(n\) decreases the SE by a factor of \(1/\sqrt{2}\), not \(1/2\).
  8. (a) False. A confidence interval estimates the population proportion, not the sample proportion. (b) True. A 95% CI is approximately \(82\% \pm 2\%\). (c) True. This is what the confidence level means: about 95% of intervals constructed this way in repeated sampling will capture \(p\). (d) True. Quadrupling \(n\) shrinks the margin of error by a factor of \(1/\sqrt{4} = 1/2\). (e) True. The 95% CI lies entirely above 50%.
  9. With a random sample, independence is satisfied and the success-failure condition holds. The margin of error is \(ME = z^{\star}\sqrt{\hat{p}(1-\hat{p})/n} = 1.96 \sqrt{(0.56)(0.44)/600} = 0.0397 \approx 4\%\).
  10. (a) No. The sample is limited to students who took the SAT and who completed an optional web survey, so it does not represent all high-school seniors. (b) \((0.5289, 0.5711)\). We are 90% confident that between 53% and 57% of SAT-taking high-school seniors are fairly certain they will study abroad. (c) In repeated random sampling, about 90% of such 90% CIs would capture the true proportion. (d) Yes — the interval lies entirely above 50%.
  11. (a) We test \(H_0: p = 0.5\) vs. \(H_A: p \ne 0.5\), with \(\hat{p} = 0.55\) and \(n = 617\). Independence holds (random sample); the success-failure condition holds using \(p_0 = 0.5\) (both \(np_0 = 308.5\) are well above 10). \(SE = \sqrt{p_0(1-p_0)/n} \approx 0.02\), so \(Z = (0.55 - 0.5)/0.02 = 2.5\). Two-tailed p-value \(= 2(0.0062) = 0.0124 < 0.05\), so we reject \(H_0\). Because \(\hat{p} > 0.5\), we have strong evidence that a majority of Independents support the plan. (b) Generally a hypothesis test and confidence interval agree, so we would expect a 95% CI entirely above 0.5. If the CI’s confidence level does not match \(1 - \alpha\) (e.g., a 99% CI paired with an \(\alpha = 0.05\) test), the CI can include 0.5 while the test still rejects.
  12. (a) The sample represents chips manufactured during that week; we should be cautious about generalizing to other weeks, since defect rates can drift over time. (b) The parameter is the true fraction of chips produced that week with severe defects. (c) \(\hat{p} = 27/212 = 0.127\). (d) The standard error (\(SE\)). (e) \(SE \approx \sqrt{\hat{p}(1-\hat{p})/n} = \sqrt{(0.127)(0.873)/212} \approx 0.023\). (f) \(\hat{p} = 0.10\) is only about one SE away from 0.127, which is not surprising; the engineer should not be alarmed. (g) Using \(p = 0.10\) instead of \(\hat{p}\): \(SE = \sqrt{(0.10)(0.90)/212} \approx 0.021\) — almost the same, as is typical when the two proportions being plugged in are similar.
  13. (a) The visitors are a simple random sample, so independence holds. Both \(64\) and \(752 - 64 = 688\) exceed 10, so the success-failure condition holds. A normal-based CI is appropriate. (b) \(\hat{p} = 64/752 = 0.085\); \(SE = \sqrt{(0.085)(0.915)/752} \approx 0.010\). (c) Using \(z^\star = 1.65\): \(0.085 \pm 1.65(0.010) \to (0.0685, 0.1015)\). We are 90% confident that between 6.85% and 10.15% of first-time visitors will register under the new design.
  14. (a) The parameter is \(p_{\text{Asian-Indian}} - p_{\text{Chinese}}\). The statistic is \(\hat{p}_{\text{Asian-Indian}} - \hat{p}_{\text{Chinese}} = 223/4373 - 279/4736 \approx -0.008\). (b) The standard error of the difference is roughly 0.005. (c) \(H_0: p_{\text{Asian-Indian}} - p_{\text{Chinese}} = 0\) vs. \(H_A: \ne 0\). The evidence is borderline; there is not strong evidence of a difference in the current-smoker rate between these two ethnic groups.
  15. (a) The standard error of \(\hat{p}_{\text{Filipino}} - \hat{p}_{\text{Chinese}}\) is roughly 0.00625. (b) We are 95% confident that the true proportion of Filipino-American current smokers is between 5.28 and 7.72 percentage points higher than that of Chinese Americans. (c) With small changes in rounding, the interval is essentially the same: 5.2 to 7.7 percentage points higher.
  16. (a) The two graphs have similar standard errors (about 0.012), but different centers: Method A is centered at 0.07 (the observed difference in sample proportions), while Method B is centered at 0. Method A is a bootstrap distribution (used for a CI), and Method B is a null distribution (used for a hypothesis test). (b) A confidence-interval question: what is the difference in the proportions of Bachelor’s vs. Associate’s students who believe COVID-19 will negatively affect their degree completion? (c) A hypothesis-test question: does the proportion of Bachelor’s students who believe COVID-19 will negatively affect degree completion differ from the proportion of Associate’s students?
  17. (a) Convert to counts: Nevirapine has 26 Yes / 94 No; Lopinavir has 10 Yes / 110 No. (b) \(H_0: p_N = p_L\) (no difference in virologic-failure rates); \(H_A: p_N \ne p_L\). (c) Random assignment gives independence within groups; the study population may not generalize beyond similar patients. The pooled proportion is \(\hat{p}_{\text{pool}} = 36/240 = 0.15\), and the success-failure condition is met using this pooled estimate. \(Z \approx 2.89\) gives \(p\text{-value} \approx 0.0039 < 0.05\), so we reject \(H_0\): there is strong evidence of a difference in failure rates between the two treatments.
  18. (a) \(SE = \sqrt{(0.79)(0.21)/347 + (0.55)(0.45)/617} \approx 0.03\). Using \(z^\star = 1.96\): \((0.79 - 0.55) \pm 1.96(0.03) \to (0.181, 0.299)\). We are 95% confident that the proportion of Democrats supporting the plan is between 18.1 and 29.9 percentage points higher than the proportion of Independents. (b) True.
  19. (a) We treat each job as a Bernoulli trial for “men paid more than women”, with \(H_0: p = 0.5\) and \(H_A: p \ne 0.5\). (b) Independence is not strictly checkable since the 21 jobs are not a random sample, but the jobs seem to be distinct enough that independence is not unreasonable. The success-failure condition uses \(p_0 = 0.5\): both \(np_0 = 21(0.5) = 10.5\) exceed 10. Compute \(\hat{p} = 19/21 = 0.905\) and \(SE = \sqrt{(0.5)(0.5)/21} \approx 0.109\), giving \(Z = (0.905 - 0.5)/0.109 \approx 3.72\). The two-tailed p-value is about \(2(0.0001) = 0.0002 < 0.05\), so we reject \(H_0\): men are paid more in a discernibly higher fraction of jobs than would be expected by chance.
  20. Conditions fail. Only 4 of 50 in the treatment group and 3 of 25 in the control group yawned — far short of 10 successes and 10 failures in each group. Therefore the sampling distribution of \(\hat{p}_T - \hat{p}_C\) is not approximately normal, and a normal-based confidence interval is not appropriate. A simulation-based (randomization or bootstrap) interval would be the honest choice here.
  21. (a) False. The confidence interval includes 0, so we cannot conclude the two proportions differ. (b) False. The correct interpretation is that we are 95% confident that Americans making less than $40,000 are between 16 percentage points less and 2 percentage points more likely to be “not at all personally affected” than those making $40,000 or more. (c) False. A lower confidence level produces a narrower interval. (d) True.
  22. (a) Type I. (b) Type II. (c) Type II.
  23. No — the two samples are not independent. The beginning-of-semester and end-of-semester surveys are taken on the same students, so the responses are paired.
  24. (a) With a normal centered at \(-0.1\) and \(SD = 0.15\), \(P(Z < -2 \times SE) = P(Z < -0.30 \text{ on original scale}) \approx 0.09\). (b) With center \(-0.4\) and \(SD = 0.145\), the area below \(-2 \times SE = -0.29\) is \(\approx 0.77\). (c) With center \(-0.1\) and \(SD = 0.0671\), the area below \(-2 \times SE\) is \(\approx 0.07\). (d) With center \(-0.4\) and \(SD = 0.0678\), the area below \(-2 \times SE\) is \(\approx 1.00\). (e) The larger the true effect \(\delta\) and the larger the sample size (which shrinks the SE), the more likely a future study is to reject the null hypothesis — i.e., the higher the power.

StatLens Exercises

  1. (a) The parameter is \(p\), the true proportion of all students at this campus who skip breakfast. (\(\hat{p} = 99/150 = 0.66\) is the sample estimate.) (b) \(H_0: p = 0.60\) vs. \(H_A: p > 0.60\)one-sided, because the claim gives a direction (“more than 60%”) stated before the data were collected. (c) Using \(p_0 = 0.60\): \(np_0 = 150(0.60) = 90\) and \(n(1-p_0) = 150(0.40) = 60\). Both are far above 10, so the success-failure condition is met and the normal-approximation \(z\)-procedure is trustworthy.
  2. (a) \(H_0: p = 0.10\) vs. \(H_A: p < 0.10\) — the claim is a lower complication rate. (b) With \(n = 62\) and \(p_0 = 0.10\): \(np_0 = 6.2\), which is below 10. The success-failure condition fails. (c) The normal-approximation \(z\)-test is not trustworthy here. This is exactly why we used a bootstrap interval for this same dataset earlier: when the counts are too small for the normal model, simulation is the honest tool. Sometimes the correct answer to “run the \(z\)-test” is “the conditions don’t support it.”
  3. (a) For the stent data, the two 95% intervals are very close — typically within about a percentage point at each endpoint. With \(n = 224\) and many events in each arm, the normal model fits well and both methods agree. (b) For the medical-consultant data the intervals are noticeably different, and the bootstrap distribution is visibly right-skewed (the sample proportion is small, so resamples pile up near the low end). A symmetric normal-based interval cannot capture that skew. (c) The normal approximation works when \(n\hat{p} \geq 10\) and \(n(1-\hat{p}) \geq 10\). The stent data clears this easily, so the sampling distribution of \(\hat{p}\) is nearly normal and the formula matches the simulation. The medical-consultant data has only a handful of complications (\(n\hat{p} < 10\)), so the sampling distribution is skewed — the formula’s symmetric interval is the wrong shape, while the bootstrap follows the actual (skewed) distribution. (d) The two methods aren’t disagreeing on a fact; they are two approximations, and they converge as conditions improve. A disagreement is a signal that the normal conditions have failed, not that someone made an error. When they diverge, trust the simulation: the bootstrap makes no normality assumption, so it stays honest exactly where the formula breaks down. This is the same reason the course teaches simulation before the formulas — simulation is the more general tool; the \(z\)-procedure is a convenient shortcut that applies only when conditions are met.

E.13 Chapter 13 — Confidence Intervals for Means

Solutions to the odd-numbered exercises in Chapter 13 — Confidence Intervals for Means.

  1. (a) Statistic: the average nightly sleep among the 25 New Yorkers sampled. Parameter: the average nightly sleep among all New Yorkers. (b) Statistic: the average height of the undergraduates in the sample. Parameter: the average height of all undergraduates at the two universities (or, more broadly, the population the samples represent).
  2. (a) Use the sample mean to estimate the population mean: 171.1 cm. Use the sample median to estimate the population median: 170.3 cm. (b) Use the sample standard deviation, 9.4 cm, and the sample IQR, \(177.8 - 163.8 = 14\) cm. (c) \(z_{180} = (180 - 171.1) / 9.4 \approx 0.95\) and \(z_{155} \approx -1.71\). Neither is more than two standard deviations from the mean, so neither height is unusual. (d) No. Sample statistics estimate the population parameter, but they vary from one sample to the next; a new sample would give slightly different values. (e) The standard error of the mean measures how much the sample mean varies from sample to sample. Using the given sample, \(SE_{\bar{x}} = s / \sqrt{n} = 9.4 / \sqrt{507} \approx 0.42\) cm.
  3. (a) The kindergartners will have a smaller standard deviation of heights — a group of same-age children is more homogeneous in height than a mixed-age group of adults. (b) The standard error depends on the individual-level variability. For adults, \(SE_{\bar{x}} \approx 9.4 / \sqrt{100} = 0.94\) cm. For kindergartners, the SE will be smaller because their individual heights vary less to begin with.
  4. (a) \(df = 5\), \(t^{\star}_{5} = 2.02\). (b) \(df = 20\), \(t^{\star}_{20} = 2.53\). (c) \(df = 28\), \(t^{\star}_{28} = 2.05\). (d) \(df = 11\), \(t^{\star}_{11} = 3.11\).
  5. (a) SE \(\approx\) 0.1 weeks (roughly the standard deviation of the bootstrap distribution of \(\bar{x}\)). (b) A 99% percentile bootstrap CI is roughly (38.45 weeks, 38.85 weeks) — the middle 99% of the bootstrap means. (c) A 99% SE bootstrap CI is \(\bar{x} \pm 2.58 \cdot SE \approx (38.49, 38.91)\) weeks. Both intervals give essentially the same story about the true mean gestation length.
  6. The interval is symmetric about the sample mean, so \(\bar{x}\) = midpoint = \((18.985 + 21.015)/2 = 20\). The margin of error is \(ME = 21.015 - 20 = 1.015\). With \(df = 36 - 1 = 35\) and 95% confidence, \(t^{\star}_{35} \approx 2.03\). Using \(ME = t^{\star} \cdot s / \sqrt{n}\), solve for \(s\): \(s = ME \cdot \sqrt{n} / t^{\star} = 1.015 \cdot 6 / 2.03 \approx 3.0\).
  7. Because \(z^{\star}\) is smaller than \(t^{\star}_{df}\), using \(z^{\star}\) would produce a narrower interval. A narrower interval captures the true mean less often in repeated sampling, so the actual coverage probability would be lower than the stated confidence level (a nominal 95% interval built with \(z^{\star}\) actually covers less than 95% of the time). The discrepancy is largest for small \(n\) and shrinks as \(n\) grows and \(t^{\star}\) approaches \(z^{\star}\).
  8. We use a hypothesis test (two-sample \(t\)-test) to evaluate whether the data provide convincing evidence of a difference between two population means, and we use a confidence interval to estimate this difference.
  9. (a) A 90% percentile bootstrap CI takes the 5th and 95th percentiles of the bootstrap distribution — read off the histogram, roughly (0.1 m/sec, 0.35 m/sec). (b) A 90% SE bootstrap CI is \(\bar{x}_{\text{diff}} \pm 1.65 \cdot SE_{\text{boot}}\), where \(\bar{x}_{\text{diff}}\) is the observed difference in means and \(SE_{\text{boot}}\) is the standard deviation of the bootstrap distribution. Both intervals should be similar because the bootstrap distribution is roughly symmetric. Both suggest Western fence lizards run faster on average than Sagebrush lizards, since the intervals lie entirely above 0.
  10. Using the observed difference in sample means and the SE reported on the bootstrap distribution (\(SE_{\text{boot}} = 4.64\)), together with \(df = \min(22, 22) = 22\) and \(t^{\star}_{22} \approx 2.07\), the 95% mathematical CI is approximately \(-12.03 \pm 2.07 \cdot 4.64 \to\) ($-2.96, $-22.42$ per carat) — that is, we are 95% confident the population average price per carat of 0.99-carat diamonds is $2.96 to $22.42 lower than that of 1-carat diamonds. (This closely matches the bootstrap intervals in Exercise 16.)
  11. (a) \(\mu_{\bar{x}_1} = 15\), \(\sigma_{\bar{x}_1} = 20 / \sqrt{50} \approx 2.83\). (b) \(\mu_{\bar{x}_2} = 20\), \(\sigma_{\bar{x}_2} = 10 / \sqrt{30} \approx 1.83\). (c) \(\mu_{\bar{x}_2 - \bar{x}_1} = 20 - 15 = 5\); \(\sigma_{\bar{x}_2 - \bar{x}_1} = \sqrt{(20/\sqrt{50})^2 + (10/\sqrt{30})^2} \approx 3.37\). (d) Thinking of \(\bar{x}_1\) and \(\bar{x}_2\) as independent random variables, the SD of their difference is the square root of the sum of their squared SDs: \(SD_{\bar{x}_2 - \bar{x}_1} = \sqrt{SD_{\bar{x}_2}^2 + SD_{\bar{x}_1}^2}\).
  12. (a) True — a paired analysis reduces to a one-sample analysis on the differences. (b) True. Paired data require one observation per unit in each group, so the two groups must have the same size. (Same size alone is not enough for pairing, though.) (c) True — that natural one-to-one correspondence is the definition of paired data. (d) False. In a paired analysis, we compute the difference between each matched pair of observations, not the difference between an observation and the other dataset’s average.
  13. (a) Paired — Intel’s and Southwest’s stock prices are recorded on the same 50 days, so each day pairs one Intel price with one Southwest price. (b) Paired — each of the 50 items has a Target price and a Walmart price for that same item. (c) Not paired — two independent random samples of 100 students are drawn from two different high schools, with no natural correspondence between individual students in one sample and individual students in the other.
  14. (a) The bootstrap distribution of the mean difference (read - write) is centered near \(\bar{x}_{r-w} = -0.545\). A 95% percentile CI takes the middle 95% of the histogram, giving approximately \((-1.79, 0.69)\) points. (b) The bootstrap SE (standard deviation of the bootstrap distribution) is roughly \(s / \sqrt{n} = 8.887 / \sqrt{200} \approx 0.63\). A 95% SE CI is \(-0.545 \pm 2 \cdot 0.63 \to (-1.81, 0.71)\). (c) We are 95% confident that in the population of students who take the HSB survey, the true average difference between reading and writing scores (read \(-\) write) is somewhere between about \(-1.8\) and \(0.7\) points. (d) No — both intervals contain 0, so the data do not provide convincing evidence of a real difference in average reading and writing scores.
  15. (a) With \(n = 200\), \(df = 199\) and \(t^{\star}_{199} \approx 1.97\) for 95% confidence. \(SE = 8.887 / \sqrt{200} \approx 0.628\). CI: \(-0.545 \pm 1.97 \cdot 0.628 \to (-1.78, 0.69)\) points. (b) We are 95% confident that the true average difference between reading and writing scores (read \(-\) write) in the population is between \(-1.78\) and \(0.69\) points. (c) No — 0 lies inside the interval, so the data do not provide convincing evidence of a real average difference between the two scores.
  16. With \(\bar{x}_{\text{diff}} = 12.5\), \(s_{\text{diff}} = 7.2\), and \(n = 50\), we have \(df = 49\) and \(t^{\star}_{49} \approx 2.68\) for 99% confidence. \(SE = 7.2 / \sqrt{50} \approx 1.02\). The 99% CI is \(12.5 \pm 2.68 \cdot 1.02 \to (9.77, 15.23)\) feet. We are 99% confident that the average growth of these young trees between 2009 and 2019 was between 9.77 and 15.23 feet. Because the interval lies entirely above 0, the data provide strong evidence of positive average growth over the decade.

StatLens Exercises

  1. (a) The parameter is \(\mu\), the population mean daily caffeine intake (mg) among all undergraduates at the university. The sample mean \(\bar{x}\) from the 48 students is the estimate, not the parameter itself. (b) Two-sided. A confidence interval is two-sided by construction — it estimates the parameter with a margin of error in both directions. One-sided procedures live on the hypothesis-testing side, for directional claims. (c) We rarely (essentially never) know \(\sigma\) in practice. The \(t\)-procedure plugs in \(s\) as the best available estimate of \(\sigma\) and widens the interval slightly (using \(t^{\star}\) instead of \(z^{\star}\)) to account for the extra uncertainty introduced by that estimate. The smaller the sample, the heavier the \(t\) tails and the wider the interval.
  2. (a) The one-sample \(t\)-interval requires (i) independence — a random sample from the population (or the result of random assignment in an experiment) — so one observation’s value says nothing about another’s; and (ii) approximate normality — either the individual observations look roughly bell-shaped and outlier-free (for small \(n\)), or the sample is large enough (typically \(n \ge 30\)) for the CLT to smooth things out. (b) At \(n = 12\) the CLT is not in play, so inspect the data directly with a dotplot or histogram. Look for strong skew, gaps, or outliers — any of those would make the \(t\)-interval untrustworthy. (c) No. A strong skew or an outlier at \(n = 12\) can pull \(\bar{x}\) and inflate \(s\), distorting the interval. Switch to a bootstrap CI for the mean (or investigate whether the outlier is a data-quality issue before proceeding). When small samples and non-normal data collide, simulation is the honest choice.
  3. (a) Each runner is their own control, so the difference \(D_i = \text{after}_i - \text{before}_i\) subtracts out the huge runner-to-runner variability in baseline ability — variability that would otherwise dominate the standard error. Only the within-runner change (typically much smaller) remains. (b) The paired interval uses \(s_D / \sqrt{n}\), where \(s_D\) is the SD of the within-runner differences — typically small. The two-sample interval ignores the pairing and pools SDs across runners — typically much larger. Same numbers, but the two-sample interval is substantially wider and may include 0 when the paired interval excludes it. (c) Compare the mean change in 5K time of the trained runners to the mean change (or, if no baseline was recorded for the untrained group, the mean final time) of the untrained runners — a two-sample \(t\). Matching is impossible because the untrained runners are different people; there is no natural correspondence. (d) The two procedures answer different questions. Paired analyzes one variable (the within-pair difference); two-sample analyzes two variables (the two group means). Both use the \(t\) distribution because both involve a sample mean with an estimated standard error, but the standard errors are computed from different things — and calling paired “just a special case” is exactly the confusion that leads people to throw away the pairing (and the power that comes with it).

E.14 Chapter 14 — Hypothesis Tests for Means

Solutions to the odd-numbered exercises in Chapter 14 — Hypothesis Tests for Means.

  1. (a) \(df = 10\), two-tailed \(p = 2 \cdot P(T_{10} \ge 1.91) \approx 0.085\). Do not reject \(H_0\). (b) \(df = 16\), two-tailed \(p = 2 \cdot P(T_{16} \ge 3.45) \approx 0.003\). Reject \(H_0\). (c) \(df = 6\), two-tailed \(p = 2 \cdot P(T_6 \ge 0.83) \approx 0.438\). Do not reject \(H_0\). (d) \(df = 27\), two-tailed \(p = 2 \cdot P(T_{27} \ge 2.13) \approx 0.042\). Reject \(H_0\).
  2. (a) Point estimate of the average pregnancy length is the sample mean; point estimate of the median is the sample median. (b) \(H_0: \mu = 40\) weeks vs. \(H_A: \mu \ne 40\) weeks. With \(n = 1000\) the sample mean is well below 40 (roughly 38.7 weeks) and \(SE = s/\sqrt{n}\) is small (\(\approx 0.08\)), so \(T\) is many standard errors below 0 and the p-value is essentially 0. Reject \(H_0\): the average length of gestation in this population is not 40 weeks. (c) The data alone cannot distinguish the two friends’ claims. Both a genuine recent decrease in gestation length and a difference in how “40 weeks” is defined would produce a sample mean below 40 in this dataset — we would need historical comparison data (to see a trend) or data on measurement conventions to tell the two explanations apart.
  3. (a) \(H_0: \mu = 8\) (New Yorkers sleep 8 hours per night on average) vs. \(H_A: \mu \ne 8\) (New Yorkers sleep less or more than 8 hours on average). (b) The sample is random, and the reported min and max suggest no extreme outliers, so conditions are reasonable. \(SE = 0.77/\sqrt{25} = 0.154\), \(T = (7.73 - 8)/0.154 = -1.75\), \(df = 24\). (c) Two-tailed p-value \(\approx 0.093\). If in fact New Yorkers averaged 8 hours of sleep per night, we would see a random sample of 25 with a mean at least this far from 8 about 9.3% of the time. (d) Since \(p > 0.05\), do not reject \(H_0\). The data do not provide strong evidence that New Yorkers sleep more or less than 8 hours per night on average. (e) Yes — since we did not reject \(H_0\), a 90% CI would include 8.
  4. (a) One-sample \(t\)-test. \(H_0: \mu = 5\) vs. \(H_A: \mu \ne 5\) at \(\alpha = 0.05\). Random sample, so observations are independent; assume the distribution of years of piano lessons is approximately normal. \(SE = 2.2/\sqrt{20} = 0.492\), \(T = (4.6 - 5)/0.492 = -0.81\), \(df = 19\). The one-tail area is about 0.21, so the p-value is about 0.42. Since \(p > 0.05\), do not reject \(H_0\): the data do not provide sufficient evidence that the average number of years of piano lessons differs from 5. (b) Using \(SE = 0.492\) and \(t^{\star}_{19} = 2.09\), the 95% CI is \(4.6 \pm 2.09 \times 0.492 = (3.57,\, 5.63)\). We are 95% confident that the average number of years students in this city take piano lessons is between 3.57 and 5.63. (c) They agree: we failed to reject \(H_0\), and the null value \(\mu = 5\) lies inside the 95% CI.
  5. The hypotheses should be about the population means, not the sample means. Use \(\mu_{AD}\) and \(\mu_I\) (not \(\bar{x}_{AD}\) and \(\bar{x}_I\)). The null should be an equality: \(H_0: \mu_{AD} = \mu_I\). And since the baker wants to know whether there is any difference, the alternative should be two-sided: \(H_A: \mu_{AD} \ne \mu_I\).
  6. \(H_0: \mu_{WF} = \mu_S\) vs. \(H_A: \mu_{WF} \ne \mu_S\), where \(\mu\) is the true mean top speed of each species. The observed difference \(\bar{x}_{WF} - \bar{x}_S = 0.7\) m/sec sits far in the right tail of the 1,000 randomized differences (which are centered near 0 and rarely reach \(\pm 0.5\)). Very few permutations produce a difference of 0.7 or more extreme, so the p-value is very small (well under 0.05). Reject \(H_0\): the data provide strong evidence that the two species differ in average top speed, with the Western fence lizard faster than the Sagebrush lizard on average.
  7. \(H_0: \mu_{0.99} = \mu_1\) vs. \(H_A: \mu_{0.99} \ne \mu_1\). Independence: both samples are random and much smaller than 10% of their respective populations, so within-sample and between-sample independence are reasonable. Normality: the box plots are not extremely skewed and show no extreme outliers, so the sampling distribution of \(\bar{x}_{0.99} - \bar{x}_1\) is approximately normal. \(T_{22} \approx -2.7\), two-tailed p-value \(\approx 0.013\). Since \(p < 0.05\), reject \(H_0\): the data provide convincing evidence that the average price per carat differs between 0.99-carat and 1-carat diamonds (with 1-carat diamonds priced higher per carat).
  8. \(H_0: \mu_T = \mu_C\) vs. \(H_A: \mu_T \ne \mu_C\), where \(\mu_T\) and \(\mu_C\) are the mean emotional-exhaustion scores in the treatment and control populations. Conditions are given to hold. \(SE = \sqrt{\tfrac{4.44^2}{30} + \tfrac{8.87^2}{30}} = \sqrt{0.66 + 2.62} = 1.81\). \(T = (15.47 - 32.43)/1.81 \approx -9.4\), \(df \approx 29\) (Welch), so the p-value is essentially 0. Reject \(H_0\): the data provide overwhelming evidence that mean emotional exhaustion is lower in the MBI treatment group than in the control group. Because this was a randomized experiment, the reduction can be attributed to the mindfulness-based intervention.
  9. \(H_0: \mu_M = \mu_A\) vs. \(H_A: \mu_M \ne \mu_A\), where \(\mu_M\) and \(\mu_A\) are the mean city fuel efficiencies (MPG) of manual and automatic 2021 cars. Both samples are random with \(n = 25\); check the box plots for extreme skew or outliers. Compute \(SE = \sqrt{s_M^2/25 + s_A^2/25}\) from the table, then \(T = (\bar{x}_M - \bar{x}_A)/SE\) with \(df \approx \min(24, 24) = 24\), and the two-tailed p-value. In this sample the manual-transmission cars have higher mean city MPG than the automatics; the resulting test statistic gives \(|T| > 2\) and a p-value below 0.05. Reject \(H_0\): the data provide convincing evidence of a difference between the mean city fuel efficiency of manual and automatic 2021 cars.
  10. \(H_0: \mu_M = \mu_A\) vs. \(H_A: \mu_M \ne \mu_A\), where \(\mu_M\) and \(\mu_A\) are the mean highway fuel efficiencies (MPG) of manual and automatic 2021 cars. The same conditions and formulas apply as in the city-MPG problem: \(SE = \sqrt{s_M^2/25 + s_A^2/25}\), \(T = (\bar{x}_M - \bar{x}_A)/SE\) with \(df \approx 24\). In this sample manual cars again show a higher mean highway MPG than automatic cars, and the corresponding p-value is small enough to reject \(H_0\) at \(\alpha = 0.05\). The data provide convincing evidence of a difference between the mean highway fuel efficiency of manual and automatic 2021 cars.
  11. \(H_0: \mu_T = \mu_C\) vs. \(H_A: \mu_T \ne \mu_C\), where \(\mu_T\) and \(\mu_C\) are the mean number of items recalled in the distracted (treatment) and control populations. \(SE = \sqrt{\tfrac{1.8^2}{22} + \tfrac{1.8^2}{22}} = \sqrt{0.147 + 0.147} = 0.543\). \(T = (4.9 - 6.1)/0.543 \approx -2.21\), \(df \approx 21\), two-tailed p-value \(\approx 0.038\). Since \(p < 0.05\), reject \(H_0\): the data provide convincing evidence that distracted eaters recall fewer lunch items on average than controls. Because subjects were randomized to conditions, the reduction can be attributed to the distraction.
  12. (a) Let \(\text{diff} = \text{temp}_{2022} - \text{temp}_{1950}\). \(H_0: \mu_{\text{diff}} = 0\) vs. \(H_A: \mu_{\text{diff}} \ne 0\). (b) The observed mean difference \(\bar{x}_{\text{diff}} = 2.53\)°F falls well outside the bulk of the 1,000 randomized differences (which are centered at 0 and rarely exceed \(\pm 1.5\)°F). (c) Almost no randomized differences are as extreme as the observed one, so the p-value is essentially 0. Reject \(H_0\): the data provide strong evidence that the 90th percentile high temperature is different in 2022 than in 1950 (higher in 2022 at the sampled NOAA stations).
  13. (a) Paired: for each station the 2022 measurement corresponds to exactly one 1950 measurement at the same location, so the two years’ observations are not independent. Analyze the 26 differences. (b) \(H_0: \mu_{\text{diff}} = 0\) (no change in 90th percentile high temperature at NOAA stations between 1950 and 2022) vs. \(H_A: \mu_{\text{diff}} \ne 0\) (there is a change). (c) The stations are described as representative of the lower 48 states, so independence between stations is reasonable. \(n = 26\) is close to 30 and the histogram shows no particularly extreme outliers, so the sampling distribution of \(\bar{x}_{\text{diff}}\) is approximately normal. (d) \(SE = 2.95/\sqrt{26} = 0.579\), \(T = (2.53 - 0)/0.579 \approx 4.37\), \(df = 25\), two-tailed p-value \(\approx 0.0002\). (e) Since \(p < 0.05\), reject \(H_0\): the data provide strong evidence that the 90th percentile high temperature at NOAA stations is different — specifically, hotter — in 2022 than in 1950. (f) A Type I error is possible: we would have rejected \(H_0\) when in fact there was no true change in average 90th percentile high temperature. (g) No — since we rejected \(H_0\) with a null value of 0, a corresponding confidence interval would lie entirely above 0.
  14. (a) Paired design: each student takes both exams; randomly assign, for each student separately, which exam is taken with music and which in silence. Analyze the within-student differences in exam scores. (b) Independent-samples design: each student is randomly assigned to use one study environment for both exams. Compare the mean exam scores across the two groups of students.
  15. (a) \(H_0: \mu_{\text{diff}} = 0\) vs. \(H_A: \mu_{\text{diff}} \ne 0\), where \(\text{diff} = \text{admissions}_{6\text{th}} - \text{admissions}_{13\text{th}}\). Using the reported sample statistics (with the paired sample sizes and standard deviation of the differences), \(T \approx -2.71\), \(df = 5\), two-tailed p-value \(\approx 0.042\). Since \(p < 0.05\), reject \(H_0\): the data provide evidence of a difference in average ER admissions between the two Fridays, with admissions higher on Friday the 13th on average. (b) The corresponding 95% CI for \(\mu_{\text{diff}}\) is approximately \((-6.49,\, -0.17)\) admissions. (c) The reported conclusion overreaches. This is an observational study, so we cannot conclude that Friday the 13th causes additional admissions — there may be confounding factors, and the sample size is very small (\(n = 6\) pairs). The observed difference is discernible, but we should describe it as an association, not evidence for the “stay home” recommendation.

StatLens Exercises

  1. (a) The parameter is \(\mu\), the population mean calorie content of all bars produced by this manufacturer. The sample mean \(\bar{x}\) from the 36 measured bars is the estimate of \(\mu\), not the parameter itself. (b) \(H_0: \mu = 220\) vs. \(H_A: \mu \ne 220\)two-sided, because the consumer group is asking whether the mean differs from 220, with no direction specified in advance. (c) If the group suspects under-filling, the alternative becomes \(H_A: \mu < 220\) (one-sided). The practical danger of switching to a one-sided test after looking at the data is that you double the effective Type I error rate: a one-sided test at \(\alpha = 0.05\) has a 5% false-alarm rate only when the direction was chosen before the data — otherwise both tails are effectively “in play” and you should report the two-sided p-value.
  2. (a) (i) Independence of observations (random sample or random assignment). (ii) Approximate normality of the data for small \(n\), or of the sampling distribution of \(\bar{x}\) (the CLT covers you for larger \(n\), where “larger” depends on how skewed the population is). (b) At \(n = 15\) the CLT does little work, so look at the picture itself: roughly symmetric, no heavy tails, no outliers? Skew, gaps, or outliers are warning signs that the \(t\) distribution won’t accurately model \(\bar{x}\). (c) Ask two questions: (i) Is the outlier a data-quality error? (typo, sensor failure, wrong unit) — if so, fix or remove it and document why. (ii) Is the outlier a real observation that simply doesn’t fit a normal model? — in which case the \(t\)-test is not the right tool: switch to a randomization test (no normality assumption), or report results both with and without the outlier and be transparent about the difference.
  3. (a) For well-behaved, moderate-\(n\) data the two p-values are typically very close — often agreeing to two decimal places. The \(t\) distribution does a good job of approximating the permutation distribution because the CLT has kicked in and both groups look roughly normal. (b) In a small-\(n\) skewed/outlier case the \(t\)-test’s p-value can be misleadingly small: a single outlier pulls one group’s mean and inflates the test statistic, while the randomization distribution — built from the actual permutations of the actual data — knows there are not many ways to produce such an extreme split. The randomization p-value is typically larger (more conservative) than the \(t\) p-value in this regime. (c) The \(t\)-test’s p-value comes from a \(t\) distribution, which is the exact sampling distribution of \(\bar{x}_1 - \bar{x}_2\) only when both populations are normal. Drop normality and the \(t\) distribution becomes an approximation; with small \(n\) and skew/outliers, the approximation is poor. The randomization test makes no normality assumption — it asks only that under \(H_0\) the labels are exchangeable, an assumption satisfied by random assignment alone. (d) The two procedures aren’t “disagreeing about a fact”; they are two procedures with different assumptions. When they diverge, that is a signal that the \(t\)-test’s assumptions have failed. Trust the randomization test in that situation: it stays honest exactly where the \(t\) distribution breaks down.

E.15 Chapter 15 — Chi-Square Goodness of Fit

Solutions to the odd-numbered exercises in Chapter 15 — Chi-Square Goodness of Fit.

  1. (a) False. The chi-square distribution has a single parameter, the degrees of freedom. (b) True. (c) True — it is a sum of squared standardized deviations, so it cannot be negative. (d) False. As the degrees of freedom increases, the distribution becomes less skewed and more symmetric (approaching normal in shape).
  2. (a) \(H_0\): \(p_\text{hard} = 0.60\), \(p_\text{print} = 0.25\), \(p_\text{online} = 0.15\). \(H_A\): at least one of these proportions differs from the claimed value. (b) Expected counts: hard copy \(= 126 \times 0.60 = 75.6\), print \(= 126 \times 0.25 = 31.5\), online \(= 126 \times 0.15 = 18.9\). (c) Independence (a random sample of the professor’s students, and each student’s format choice is independent of the others) and expected counts \(\geq 5\) in every category (75.6, 31.5, 18.9 — all satisfied). (d) \(\chi^2 = \dfrac{(71-75.6)^2}{75.6} + \dfrac{(30-31.5)^2}{31.5} + \dfrac{(25-18.9)^2}{18.9} \approx 0.28 + 0.07 + 1.97 = 2.32\), with \(df = 3 - 1 = 2\), giving p-value \(\approx 0.31\). (e) Since \(p \approx 0.31 > 0.05\), fail to reject \(H_0\). The data do not provide convincing evidence that the professor’s predicted distribution of book formats is inaccurate.

StatLens Exercises

  1. (a) A distribution across categories: the parameters are the five color proportions \((p_R, p_O, p_Y, p_G, p_B)\), which must sum to 1. (b) \(H_0\): the true distribution matches the claim, \((p_R, p_O, p_Y, p_G, p_B) = (0.30, 0.20, 0.20, 0.15, 0.15)\). \(H_A\): at least one of these proportions differs from the claimed value. (c) The \(\chi^2\) statistic sums squared deviations between observed and expected counts, so it cannot be negative — any departure from \(H_0\), in any category and in any direction, inflates it. There is no meaningful one-sided GOF because there is no single direction in which the data could disagree with a full distributional claim.
  2. (a) The expected count in every category should be at least 5. We use expected rather than observed counts because the test relies on a normal approximation to the sampling variability of the counts under \(H_0\); when the expected counts are too small, that approximation — and the \(\chi^2\) reference distribution built from it — becomes unreliable regardless of what the observed counts happen to be. (b) The \(\chi^2\) p-value is not trustworthy. Use a simulation approach instead — StatLens’s simulate/goodness-of-fit/ draws repeated multinomial samples under \(H_0\) to build the actual null distribution of the \(\chi^2\) statistic, making no large-sample assumption. The underlying issue is that with small expected counts the test statistic has a discrete distribution that does not match the smooth \(\chi^2\) curve, especially in the tails. (c) Yes, you can merge a small-expected category with a substantively similar neighbor. Gained: larger expected counts, restoring the validity of the \(\chi^2\) approximation. Lost: the ability to say which of the merged categories was responsible for any departure from \(H_0\). Merging just to inflate counts (without a substantive rationale for the merge) is data-massaging.
  3. (a) The simulated null distribution of \(\chi^2\) is right-skewed, supported on \([0, \infty)\), with the observed \(\chi^2\) marked as a vertical reference on the dotplot. Its shape approximates a \(\chi^2\) distribution with the same degrees of freedom — when the expected counts are large. (b) For well-behaved data the simulation and the approximation agree very closely, often to two decimal places. This confirms that the approximation is doing what it is supposed to. (c) With small expected counts the simulation p-value is the trustworthy one. The \(\chi^2\) approximation breaks down precisely in this regime; the simulation makes no large-sample assumption and just builds the actual null distribution of the test statistic. (d) The framing is backwards. The simulation is the gold standard, and the formula is an approximation to it that happens to be convenient when the conditions are met. Agreement is confirmation that the formula’s conditions hold, not redundancy; disagreement is a signal that the formula’s conditions have failed, and the simulation is the one to trust. This is exactly the same logic as bootstrap-vs-\(t\) in the earlier CI chapters.

E.16 Chapter 16 — Chi-Square Test of Independence

Solutions to the odd-numbered exercises in Chapter 16 — Chi-Square Test of Independence.

  1. (a) The two-way table appears below. (b-i) \(E_{\text{patch}+\text{group},\,\text{quit}} = \dfrac{(150)(70)}{300} = 35\); the observed count (40) is higher than expected. (b-ii) \(E_{\text{only patch},\,\text{did not quit}} = \dfrac{(150)(230)}{300} = 115\); the observed count (120) is higher than expected.

    | Treatment              | Quit: Yes | Quit: No | Total |
    |------------------------|----------:|---------:|------:|
    | Patch + support group  |        40 |      110 |   150 |
    | Only patch             |        30 |      120 |   150 |
    | Total                  |        70 |      230 |   300 |</li>
  2. (a) Column proportions: Sun \(\approx 0.343\), Partial \(\approx 0.325\), Shade \(\approx 0.331\). (b) Expected counts by habitat (order: sun, partial, shade). Desert: 40.9, 38.7, 39.4. Mountain: 36.7, 34.8, 35.5. Valley: 36.4, 34.5, 35.1. (c) Yes — the observed proportions of the three sunlight categories differ across habitats. (d) A visible difference is not a formal conclusion; we need a \(\chi^2\) test to decide whether the differences exceed what shuffling alone would produce.
  3. The observed dataset should produce a larger \(\chi^2\) statistic than a single random shuffle, because the shuffle destroys any real association between habitat and sunlight preference while the observed data retains it.
  4. (a) \(H_0\): habitat and sunlight preference are independent. \(H_A\): habitat and sunlight preference are associated. (b) After 1000 shuffles the randomization \(\chi^2\) values range from 0 up to roughly 15, and the observed \(\chi^2\) sits far to the right of that bulk. (c) The p-value is essentially 0, so we reject \(H_0\). There is strong evidence that habitat and sunlight preference are associated — knowing the habitat tells us something about how lizards choose sun, partial, or shade.
  5. (a) \(H_0\): habitat and sunlight preference are independent. \(H_A\): they are associated. (b) With the larger sample the randomization \(\chi^2\) values range roughly from 0 to 25, and the observed statistic again falls far above the null bulk. (c) The p-value is approximately 0; reject \(H_0\). There is convincing evidence of an association between habitat and sunlight preference. (d) Larger samples increase power — the probability of detecting a real association when one exists — so a real effect that was borderline with the smaller sample now shows up decisively.
  6. (a) False. The \(\chi^2\) distribution has one parameter, degrees of freedom. (b) True. (c) True. (d) False. As df increases, the \(\chi^2\) distribution becomes more symmetric (and more bell-shaped), not less.
  7. \(H_0\): sleep-quality levels and profession are independent. \(H_A\): they are associated. Observations are independent and the expected counts are all large enough, so a \(\chi^2\) test is appropriate. The test statistic is \(\chi^2 = 1\) on 2 df, giving p-value \(\approx 0.6\). Since the p-value is well above \(\alpha = 0.05\), we fail to reject \(H_0\): the data do not provide convincing evidence of an association between sleep levels and profession.
  8. (a) \(H_0\): age of Los Angeles residents is independent of shipping-carrier preference. \(H_A\): age and shipping-carrier preference are associated. (b) The conditions are not satisfied — several expected counts fall below 5, so the \(\chi^2\) approximation is not trustworthy. A randomization \(\chi^2\) test (or collapsing sparse categories) would be more defensible.

:::

StatLens Exercises

  1. (a) The two categorical variables are major (4 levels) and caffeine source (4 levels); the contingency table is \(4 \times 4\) with 16 cells. (b) \(H_0\): major and caffeine source are independent (knowing one tells you nothing about the other). \(H_A\): major and caffeine source are associated. (c) A test of independence applies when a single random sample is drawn and both variables are measured; a test of homogeneity applies when separate random samples are drawn from each level of one variable and the second variable is measured within each sample. The arithmetic is identical but the language differs: “are these characteristics related?” vs. “do these subpopulations have the same distribution?” Sampling design dictates which framing to use.
  2. (a) Expected counts should all be at least 5. Expected counts (not observed) are the criterion because the \(\chi^2\) approximation is a large-sample argument that depends on the predicted cell sizes under \(H_0\). (b) The expected count in cell \((i, j)\) is \(\dfrac{(\text{row } i \text{ total}) \times (\text{col } j \text{ total})}{\text{grand total}}\). In words: if the variables were independent, the joint frequency would be the product of the marginal proportions; multiplying by the grand total scales that back to a count. (c) A smallest expected count of 4 is below the guideline, so trust in the \(\chi^2\) approximation is shaky. Two defensible alternatives: (i) use the randomization \(\chi^2\) test (StatLens simulate/randomization-chisq/), which builds the null distribution directly without a large-sample assumption; (ii) collapse categories that are substantively similar to boost expected counts, accepting the loss of granularity.
  3. (a) For well-behaved data the simulation p-value and the \(\chi^2\)-distribution p-value agree closely — often to two or three decimal places. (b) With a smaller table or a small cell the two diverge; the simulation is the trustworthy one because the \(\chi^2\) approximation can produce artificially small p-values in sparse tables, leading to over-rejection of \(H_0\). (c) In a \(2 \times 2\) table, the \(\chi^2\) statistic equals \(z^2\) from the corresponding two-proportion test; both procedures ask the same question (“is the success rate the same in the two groups?”) and give the same two-sided p-value, so the equality is algebraic, not coincidental. (d) The omnibus \(\chi^2\) test answers “is there any association?” but not “which row-column combinations are responsible.” Useful follow-ups: examine the standardized residuals \((O - E)/\sqrt{E}\) cell-by-cell (large absolute values flag unusual cells), or run targeted \(2 \times 2\) sub-tables on the comparisons of interest — with a Bonferroni-style adjustment to \(\alpha\) if many are tested.

E.17 Chapter 17 — Analysis of Variance

Solutions to the odd-numbered exercises in Chapter 17 — Analysis of Variance.

  1. Alternative. Large differences between group means inflate the between-group variability (MSG), pushing the \(F\)-statistic above 1 and providing evidence against \(H_0\): all population means equal.
  2. (a) The means across the original data are more variable than across the randomized data — once host species is scrambled, the group means collapse toward a common center. (b) The within-species standard deviations are about the same in both plots; permuting labels doesn’t change how spread each species’ egg lengths are, only how the group means line up. (c) The original data produces the larger \(F\)-statistic: bigger between-group variability (MSG) with roughly the same within-group variability (MSE) means a larger MSG/MSE ratio.
  3. \(H_0\): \(\mu_1 = \mu_2 = \cdots = \mu_6\) (mean chick weight is the same across all six feed supplements). \(H_A\): at least one pair of feed-type mean weights differs. Conditions. Independence: chicks were randomly assigned to feed types and (presumably) housed separately, so independence within and between groups is reasonable. Approximate normality: the within-group boxplots look fairly symmetric with no severe outliers. Constant variance: the group SDs are of comparable magnitude; small differences are plausible from sampling variability given the small \(n\) per group. Test. \(F_{5,\,65} = 15.36\) with p-value \(\approx 0\). Reject \(H_0\). The data provide convincing evidence that mean weight differs across at least some of the six feed supplements.
  4. (a) \(H_0\): mean weekly MET is the same across all five coffee-consumption levels. \(H_A\): at least one pair of coffee-level mean METs differs. (b) Independence: no collection detail is given, so we must assume the participants are independent of one another (in practice, we’d want to verify). Normality: MET is bounded below by 0 and the group SDs exceed the means, indicating strong right skew — but sample sizes are enormous (\(n\) in the thousands per group), so the CLT covers the mean. Constant variance: SDs across the five groups (21.1, 25.5, 22.5, 22.0, 22.0) are similar enough to accept this condition. (c) \(F_{4,\,50734} = 5.2\), p-value \(= 0.0003\). Reject \(H_0\). The data provide convincing evidence that mean MET differs between at least one pair of coffee-consumption groups.
  5. (a) \(H_0\): mean GPA is the same for all three major groups. \(H_A\): at least one pair of mean GPAs differs. (b) The p-value exceeds 0.05, so we fail to reject \(H_0\). The data do not provide convincing evidence that the mean GPA differs across the three major groups. (c) Residual df + group df = \(195 + 2 = 197\), so total df \(= 197\), and the sample size is \(197 + 1 = 198\) students.
  6. (a) The left randomization distribution corresponds to Dataset B. Dataset B has large within-group spread (SD \(\approx 20\)), so the observed \(F\) is small and sits inside the bulk of the null distribution — the red line lands where the shuffled \(F\)-values commonly fall. (b) The right randomization distribution corresponds to Dataset A. Dataset A has small within-group spread (SD \(\approx 5\)), which inflates the observed \(F\); the red line sits far out in the right tail relative to the shuffled null.
  7. (a) The significant \(F\)-test tells you that at least one pair of population means differs; it does not tell you which pair(s) or by how much. Follow-up pairwise comparisons are needed to see the structure. (b) Below 95%. Each individual CI misses with probability \(\approx 0.05\); with six CIs the chance that at least one misses is much higher. If the errors were independent, the joint coverage would be roughly \((0.95)^6 \approx 0.74\) — a joint miss rate near 26%. (c) Bonferroni (use \(\alpha/k\) per CI), Tukey HSD (built for all-pairs after ANOVA), or Scheffé (any linear contrasts). Covered in Ch 19. (d) In practice you act on the full set of CIs, so the trustworthiness of the set — not just each individual CI — is what matters. Six CIs each at 95% deliver a much weaker collective guarantee than the individual number suggests.

StatLens Exercises

  1. (a) Explanatory: fertilizer brand (categorical, 4 levels). Response: yield (quantitative, bushels/acre). (b) \(H_0: \mu_1 = \mu_2 = \mu_3 = \mu_4\) — all four population mean yields are equal. \(H_A:\) at least one \(\mu_i\) differs from the others (not “all four differ from each other”). (c) Comparing \(k\) means simultaneously reduces to asking whether the variability between group means is larger than the variability within groups. The \(F\)-statistic = (between-group variance) / (within-group variance). Under \(H_0\) both estimates measure the same noise, so \(F \approx 1\). Under \(H_A\) real group differences inflate the numerator, pushing \(F\) above 1 — so a right-tail \(F\)-test on the variance ratio answers the question about means.
  2. (a) Name the three conditions: (i) independence within each group (random sample or random assignment), (ii) approximate normality of the response in each group (small \(n\)) or via the CLT (larger \(n\)), and (iii) equal variances across groups (the within-group SDs are roughly the same). (b) Boxes with wildly different heights (IQRs) are the tell for a failed equal-variance condition. A common rule of thumb: if the largest group SD is more than \(\approx 2\times\) the smallest, the equal-variance assumption is a stretch. (c) With \(n_1 = 8\) and a strong right skew, group 1’s sample mean is unlikely to be near-normal (the CLT hasn’t rescued it yet), so the \(F\)-test’s p-value can be misleading. A defensible alternative is the randomization ANOVA (simulate/randomization-anova/): shuffle group labels many times to build the actual null distribution of \(F\) — no normality assumption, only exchangeability under \(H_0\).
  3. (a) The simulated null distribution of \(F\) is right-skewed, supported on \([0, \infty)\), and peaks near 1 (under \(H_0\), MSG and MSE both estimate the same noise, so their ratio \(\approx 1\)). With enough shuffles it tracks the theoretical \(F\) curve — provided the ANOVA conditions hold. (b) For well-behaved data the two p-values agree closely. (c) With \(n = 6\) and a clear outlier the two p-values can diverge, and the simulation is the safer report: the \(F\)-distribution p-value can be misleadingly small when normality or equal-variance fails, but the randomization p-value only relies on exchangeability under \(H_0\). (d) Two reasons: (i) understanding the \(F\)-distribution — why it has two df parameters, why \(F \approx 1\) under \(H_0\) — builds intuition that shuffling alone won’t; (ii) once conditions are known to hold, the \(F\)-formula returns a p-value instantly, while the simulation needs thousands of shuffles. On well-behaved data the formula is the practical choice; the simulation is the safety net.

E.18 Chapter 18 — Multiple Comparisons

Solutions to the odd-numbered exercises in Chapter 18 — Multiple Comparisons.

  1. (a) False. As the number of groups increases, so does the number of pairwise comparisons, and hence the Bonferroni-modified discernibility level \(\alpha^\star = \alpha / K\) decreases (not increases). (b) True. The residual degrees of freedom equal \(n - k\), so they grow as \(n\) grows for fixed \(k\). (c) True. ANOVA is fairly robust to modest departures from equal variance when the group sample sizes are balanced. (d) False. Independence is required regardless of the total sample size — a large \(n\) does not fix dependent observations.

StatLens Exercises

  1. (a) The per-test rate is the false-positive probability of a single test run in isolation — 0.05 by convention. The family-wise rate is the probability of at least one false positive across the whole family of tests. Per-test \(\alpha\) stays fixed; family-wise \(\alpha\) grows with the number of tests. (b) Under independence and all-true nulls, \(P(\text{at least one false positive}) = 1 - (0.95)^{10} \approx 0.401\) — about a 40% chance of a spurious “discovery” somewhere in the family, even though each individual test looks safe. (c) The decision being reported is “conditions \(A\) and \(C\) differ” — a specific pair singled out from many. If the family-wise error rate is unbounded, the study cannot credibly claim to have located “the” difference. Uncorrected pairwise tests are how false discoveries slip into multi-group studies.
  2. (a) Tukey HSD — the standard follow-up when the omnibus \(F\) rejects and you want all pairwise comparisons on the same footing; usually more powerful than Bonferroni for this exact setting. (b) No post-hoc needed — without an omnibus rejection, doing pairwise tests inflates Type I error without a family-level rejection to justify them. Report the null result. (c) Bonferroni across the 2 pre-specified comparisons. Bonferroni is more flexible than Tukey (works for any pre-specified collection), and with only 2 tests it costs little power. (d) Not a multiple-comparisons problem — two groups means one test (two-sample \(t\)), no correction needed. (e) Bonferroni (or an FDR procedure like Benjamini–Hochberg). With 100 tests the family-wise error rate at uncorrected \(\alpha = 0.05\) is enormous; genomics-scale problems typically use FDR control, but the principle is the same — correct for multiplicity.
  3. (a) Tukey gives narrower CIs on most datasets — it is designed specifically for all-pairwise comparisons of means and exploits the correlation structure of those comparisons. Bonferroni ignores that structure and simply penalizes by the number of tests. (b) With few pairs (e.g., 3 groups \(\to\) 3 pairs) the two methods are closer; Tukey still edges out, but by less. With many pairs (6 groups \(\to\) 15 pairs) Tukey pulls further ahead. (c) Bonferroni controls the family-wise error rate for any collection of tests, treating each as if independent, so the correction is generic (\(\alpha / m\)). Tukey HSD is built specifically for pairwise means from the same ANOVA; it uses the studentized-range distribution, which accounts for the correlation among sample-mean differences (\(\bar{x}_1 - \bar{x}_2\) and \(\bar{x}_1 - \bar{x}_3\) share \(\bar{x}_1\)). That extra information yields tighter intervals at the same family-wise coverage. (d) Prefer Bonferroni when: (i) you have a pre-specified set of comparisons that are not all pairwise (e.g., treatment vs. control and treatment vs. combination, but not control vs. combination); (ii) the tests mix procedure types (some means, some proportions, some correlations); or (iii) the sample sizes are wildly unbalanced — Tukey’s equal-\(n\) guarantee weakens in the unbalanced case. (e) Both methods provide honest family-wise control at the same nominal level, but they are not identical — Bonferroni’s per-pair threshold is \(\alpha / m\), while Tukey’s uses the studentized-range distribution, and the two can cross a per-pair boundary at slightly different points. When they disagree on a particular pair, they are answering slightly different questions — neither is “right.” For the standard all-pairwise-of-means question, Tukey is the default; report both if the disagreement matters to the conclusion.

E.19 Chapter 19 — Correlation

Solutions to the odd-numbered exercises in Chapter 19 — Correlation.

  1. (a) Strong relationship, but a straight line would not fit the data. (b) Strong relationship, and a linear fit would be reasonable. (c) Weak relationship, and trying a linear fit would be reasonable. (d) Moderate relationship, but a straight line would not fit the data. (e) Strong relationship, and a linear fit would be reasonable. (f) Weak relationship, and trying a linear fit would be reasonable.
  2. (a) Exam 2. The scatterplot of course grade versus Exam 2 shows less scatter, so Exam 2 is more tightly associated with the course grade. The Exam 1 relationship also looks slightly nonlinear. (b) (Answers may vary.) If Exam 2 is cumulative, it might be a better indicator of overall class performance and therefore a more useful predictor.
  3. (a) \(r = -0.7 \rightarrow (4)\). (b) \(r = 0.45 \rightarrow (3)\). (c) \(r = 0.06 \rightarrow (1)\). (d) \(r = 0.92 \rightarrow (2)\).
  4. (a) There is a moderate, positive, and linear relationship between shoulder girth and height. (b) Changing units — even for just one variable — does not change the form, direction, or strength of a relationship. Correlation is unit-invariant.
  5. (a) There is a somewhat weak, positive, possibly linear relationship between distance traveled and travel time. Note the clustering near the lower-left corner. (b) Changing units does not change the form, direction, or strength of the relationship: longer distances in miles that go with longer times in minutes will still go together when the units become kilometers and hours. (c) Correlation is unaffected by unit changes, so \(r = 0.636\) in either set of units.
  6. In every part, one variable is an exact linear function of the other, so the scatterplot is a perfect line and \(r = 1\). (1) \(\text{carbs} = \text{meat} - 3\). (2) \(\text{carbs} = \text{meat} + 2\). (3) \(\text{carbs} = 2 \times \text{meat}\). The slope is positive in all three, so the correlation is exactly \(+1\) in all three. (Sketch a scatterplot with five mock countries at meat = 10, 20, 50, 75, 100 kg and compute the matching carbs values to see it directly.)
  7. (a) True. Correlation is bounded between \(-1\) and \(+1\); strength is measured by \(|r|\), so \(|-0.90| = 0.90 > 0.5\) indicates a stronger linear relationship. (b) False. Correlation measures the linear association between two numerical variables. Two categorical variables (or two numerical variables related by a curve) require a different measure of association.
  8. (a) \(r = 0.7 \rightarrow (1)\). (b) \(r = 0.09 \rightarrow (4)\). (c) \(r = -0.91 \rightarrow (2)\). (d) \(r = 0.96 \rightarrow (3)\).

StatLens Exercises

  1. (a) & (b) Answers will vary by round. A common pattern is that students over-estimate \(|r|\) on weak correlations (a hint of pattern reads as stronger than it is) and slightly under-estimate on very strong ones. (c) \(r\) summarizes only the linear component of the relationship. Two data sets can share the same \(r\) but tell very different stories — a clean line and a curve that both register the same \(r\), for instance, or an outlier that inflates \(r\) from near-zero to moderate. A perfectly determined curve like \(y = x^2\) on \([-1, 1]\) even has \(r = 0\). Always look at the scatterplot before quoting a correlation; the picture catches what the number cannot.
  2. (a) Three warning signs: (i) a non-linear pattern (curves, U-shapes) — \(r\) measures only the linear piece and will understate a real relationship; (ii) influential outliers that visibly move \(r\) when included or excluded; (iii) clusters or sub-groups within the scatter, which hint that a single \(r\) is hiding two different relationships (Simpson’s-paradox territory). (b) “There is no linear association (\(r \approx 0\)), but the scatterplot shows a clear curved relationship.” Calling this “no association” is misleading — the variables are clearly related; the linear summary simply fails to capture it. (c) Report both and discuss the influential observations openly. Removing outliers without mentioning them is data-massaging. Readers should be able to see that \(r\) is sensitive here — that sensitivity is itself part of the finding.
  3. (a) A likely confounder is temperature / summer weather: hot days drive both ice-cream sales (more buyers) and swimming activity (more opportunity for drownings). Neither variable causes the other; both respond to the same underlying driver. (b) The original \(r\) still summarizes a real observed association, but it is not a causal one. Once temperature is controlled, ice-cream sales and drownings are no longer correlated — exactly the pattern a common-cause explanation predicts. (c) A randomized controlled experiment. Random assignment breaks the confounding link by making the treatment groups similar, on average, on every potential confounder (measured or not). With random assignment, the correlation between treatment and outcome can be read causally. (d) Two reasons: (i) confounding — coffee drinkers may differ systematically from non-drinkers on age, income, lifestyle, or baseline health in ways that explain the association; (ii) reverse causation — sicker people may avoid coffee because of their symptoms, not because coffee causes better health. Either alone is enough to undermine a causal reading of an observational finding.

E.20 Chapter 20 — Linear Regression

Solutions to the odd-numbered exercises in Chapter 20 — Linear Regression.

  1. Correlation: no units. Intercept: calories (cal). Slope: cal/cm.
  2. (a) First compute the slope: \(b_1 = r \cdot s_y/s_x = 0.636 \times 113/99 \approx 0.726\). Then use the fact that the least-squares line passes through \((\bar{x}, \bar{y})\): \(b_0 = \bar{y} - b_1 \bar{x} = 129 - 0.726 \times 108 \approx 51\). Line: \(\widehat{\text{travel time}} = 51 + 0.726 \times \text{distance}\). (b) Slope: for each additional mile between stops, the model predicts an additional 0.726 minutes of travel time on average. Intercept: a distance of 0 miles corresponds to a predicted travel time of 51 minutes — not meaningful in context (zero distance means no travel); the intercept only anchors the line’s height. (c) \(R^2 = 0.636^2 \approx 0.40\). About 40% of the variability in travel time is explained by the distance between stops. (d) \(\widehat{\text{travel time}} = 51 + 0.726 \times 103 \approx 126\) minutes. (e) Residual \(= 168 - 126 = 42\) minutes: the actual travel time is 42 minutes longer than the model predicts (the model under-predicts here). (f) No — 500 miles is well outside the range of distances used to fit the model, so predicting there would be extrapolation.
  3. (a) \(\widehat{\text{poverty}} = 4.60 + 2.05 \times \text{unemployment\_rate}\). (b) The model predicts a poverty rate of 4.60% for a county with 0% unemployment. No US county has an unemployment rate that low, so this value is not directly meaningful — it just anchors the line. (c) For each 1 percentage-point increase in unemployment rate, the model predicts an average increase of 2.05 percentage points in poverty rate. (d) Unemployment rate explains about 46% of the county-to-county variability in poverty rate. (e) \(r = \sqrt{0.46} \approx 0.68\) (positive, since the slope is positive).
  4. (a) The scatterplot shows a negative association, so \(r = -\sqrt{R^2/100}\); using the printed \(R^2\), \(r \approx -\sqrt{0.72} \approx -0.85\) (negative). (b) \(b_1 = r \cdot s_y/s_x \approx -0.85 \times 16.95/26.72 \approx -0.54\); \(b_0 = \bar{y} - b_1 \bar{x} = 30.88 - (-0.54)(30.83) \approx 47.5\). Line: \(\widehat{\text{helmet}} = 47.5 - 0.54 \times \text{lunch}\). (c) Intercept: in a neighborhood where no children receive reduced-fee lunches, the model predicts about 47.5% of bike riders wear a helmet. (d) Slope: for each 1 percentage-point increase in children receiving reduced-fee lunches, the model predicts helmet use to drop by about 0.54 percentage points on average. (e) \(\widehat{\text{helmet}} = 47.5 - 0.54(40) \approx 25.9\); residual \(= 40 - 25.9 = 14.1\) percentage points. This neighborhood has helmet use about 14 percentage points higher than the model predicts for its lunch rate.
  5. (a) \(H_0: \beta_1 = 0\) (father’s age and baby’s weight are not linearly related). \(H_A: \beta_1 \neq 0\). (b) The observed slope 0.005 sits comfortably inside the bulk of the null distribution, so the p-value is large (many permuted slopes are at least as extreme). We fail to reject \(H_0\): the data do not provide convincing evidence that father’s age is associated with baby’s weight, so father’s age does not appear to be a useful predictor of baby’s weight in these data. (c) Yes — the mathematical-model p-value is also large, and both approaches lead to the same fail-to-reject conclusion.
  6. (a) Predicted weight = \(b_0 + b_1 \times 30\) (read \(b_0\) and \(b_1\) from the estimate column of the printed regression table; the result is close to 7.1 lbs, near the mean baby weight). (b) \(H_0: \beta_1 = 0\) vs. \(H_A: \beta_1 \neq 0\). The reported p-value is large (well above 0.05), so we fail to reject \(H_0\). There is no convincing evidence of a linear relationship between father’s age and baby’s weight. (c) Because we fail to reject the null of no association, father’s age is not a useful predictor of baby’s weight.
  7. (a) Read the 2.5th and 97.5th percentiles off the bootstrap histogram of slopes. The middle 95% of the bootstrap slopes runs roughly from about \(-0.003\) to about \(0.013\) lb per year of father’s age. (b) With 95% confidence, we estimate that each additional year of father’s age is associated with somewhere between a \(0.003\)-lb decrease and a \(0.013\)-lb increase in mean baby weight. Because 0 is inside the interval, the data are consistent with no linear relationship between father’s age and baby’s weight.
  8. (a) The bootstrap SE is the standard deviation of the bootstrap slopes; reading the spread of the histogram, \(SE \approx 0.004\) lb/year. (b) 95% SE interval: \(b_1 \pm 1.96 \cdot SE \approx 0.005 \pm 1.96(0.004) \approx (-0.003,\ 0.013)\). (c) With 95% confidence, we estimate the true slope — the mean change in baby’s weight (lbs) per additional year of father’s age — lies between about \(-0.003\) and \(0.013\). Since 0 falls inside the interval, the data do not rule out a slope of zero.
  9. (a) \(H_0: \beta_1 = 0\) vs. \(H_A: \beta_1 \neq 0\), where \(\beta_1\) is the population slope for annual murders per million on percent living in poverty. (b) The p-value from the regression output is very small, so we reject \(H_0\) at any conventional level. The data provide convincing evidence that poverty percentage is a useful linear predictor of annual murder rate across metropolitan areas: areas with higher poverty tend to have higher annual murder rates. (c) 95% CI: \(b_1 \pm t^*_{18} \cdot SE(b_1) = 2.559 \pm 2.101 \cdot SE(b_1)\); reading \(SE(b_1)\) from the output, the interval is roughly \((1.8,\ 3.3)\) murders per million per percentage point of poverty. (d) Yes: the 95% CI does not contain 0, and the hypothesis test rejects \(H_0\); both approaches lead to the same conclusion.
  10. (a) The bootstrap SE is the standard deviation of the 1,000 bootstrap slopes; from the spread of the histogram, \(SE \approx 0.35\) murders per million per percentage point of poverty. (b) 90% SE CI: \(2.559 \pm 1.645 \cdot 0.35 \approx (1.98,\ 3.14)\). (c) With 90% confidence, we estimate that a one-percentage-point increase in poverty is associated with between about 2.0 and 3.1 additional annual murders per million people on average. Since 0 is not in the interval, poverty percentage is a useful predictor of annual murder rate.
  11. (a) The relationship between cans of beer and BAC is positive, roughly linear, and strong (little scatter around a clear upward trend). (b) Read the coefficients from the regression table: \(\widehat{\text{BAC}} = b_0 + b_1 \cdot \text{beers}\), with \(b_0 \approx -0.013\) and \(b_1 \approx 0.018\). Slope: each additional can of beer is associated with a mean BAC increase of about 0.018 g/dL. Intercept: a person who drank 0 beers is predicted to have a slightly negative BAC — physically impossible; the intercept only anchors the fitted line and is not meaningful on its own. (c) \(H_0: \beta_1 = 0\) vs. \(H_A: \beta_1 > 0\) (or \(\neq 0\) for a two-sided test). The p-value in the regression output is very small, so we reject \(H_0\): the data provide convincing evidence that drinking more cans of beer is associated with higher BAC. (d) \(R^2 = r^2 \approx 0.89^2 \approx 0.79\): about 79% of the variability in BAC is explained by the number of cans of beer. (e) No — a bar sample would introduce many extra sources of variability the OSU experiment held roughly constant: differences in weight, food consumption, drinking pace, drink strength, time since drinking, and biological variation. Those additional sources of noise would weaken the linear relationship between number of drinks and BAC.

StatLens Exercises

  1. (a) Explanatory: square footage; response: price. The population slope is \(\beta_1\); the reported 0.15 is \(b_1\), the sample estimate. (b) \(H_0: \beta_1 = 0\) vs. \(H_A: \beta_1 \neq 0\). (c) “For each additional square foot of house, the mean price is predicted to increase by about $150.” (Convert 0.15 thousand-dollars per sqft into $150 per sqft.) The most common student errors are (i) omitting “mean” (the slope describes the average response, not what any single house will sell for), (ii) omitting units, and (iii) reversing cause and effect — the slope describes association, not “adding a square foot causes the price to rise $150.”
  2. (a) The four regression conditions are Linearity (the true relationship is approximately linear), Independence (observations are independent), Normality of residuals (residuals are approximately normal), and Equal variance of residuals (spread of residuals is constant across \(x\)). (b) Linearity fails when the residual plot shows a systematic curved pattern (a U-shape, arc, or wave) instead of random scatter around 0. Equal variance fails when the spread of residuals grows or shrinks with \(x\) — the classic fan or trumpet shape. (c) No — don’t discard automatically. First ask: (i) is it a data-quality error (typo, unit mix-up, instrument failure)? If yes, fix or remove and document the reason. (ii) Is it a real but influential observation the linear model can’t accommodate? If yes, the model is the problem, not the point — consider a transformation, a non-linear model, or transparently report the with- and without-outlier fits.
  3. (a) For a well-behaved (roughly linear, homoscedastic) dataset the two 95% intervals for \(\beta_1\) typically agree to within a small fraction of the standard error. (b) On heteroscedastic data (visible fan shape in the residuals), the bootstrap interval is typically wider than the \(t\)-interval, because it correctly accounts for the higher variability at some values of \(x\). The \(t\)-interval is artificially narrow here — the equal-variance assumption it depends on has failed. (c) The bootstrap is more honest: it does not pretend the spread is constant when it isn’t. The \(t\)-interval looks more precise, but its actual coverage falls below the advertised 95%. (d) Transforming (e.g., log–log) is the right first move when there is a theoretical reason to expect linearity on the transformed scale — but sometimes no transformation works, or the transformed slope no longer answers the original question. In those cases the bootstrap CI for the slope on the original scale is the right tool: it gives an honest interval without forcing the data into a shape they don’t have.

E.21 Chapter 21 — Prediction and Model Fit

Solutions to the odd-numbered exercises in Chapter 21 — Prediction and Model Fit.

  1. (a) The residual plot will show residuals randomly distributed around 0 with approximately constant variance. (b) The residuals will show a fan shape, with higher variability for smaller \(x\) and lower variability for larger \(x\); there will also be several points on the right sitting above the line. The linear model is not a good fit here.
  2. Over-estimate. The residual is \(e = y - \hat{y}\), so a negative residual means the observed value is smaller than the predicted value — the model predicted a longer shelf life than the apple actually had.
  3. (a) There is a positive, moderate, roughly linear association between number of calories and amount of protein. Protein is more variable for menu items with higher calorie counts, indicating non-constant variance. Two loose clusters are visible: a small group of low-calorie items in the lower left and a larger group of higher-calorie items on the right. (b) Explanatory (predictor): calories. Response (outcome): amount of protein (in grams). (c) A regression line lets us predict the amount of protein for a menu item given only its calorie count — useful when Starbucks displays calories but not the nutritional breakdown. (d) The residuals fan out with predicted value: items with higher predicted protein show more prediction error than items with lower predicted protein, so the model is more trustworthy at the low end than at the high end.
  4. (a) An outlier in the bottom right. It is far from the center of the data in the \(x\)-direction, so it has high leverage. It is also influential: without that observation, the regression line would have a very different slope. (b) An outlier in the bottom right. It has high leverage (extreme \(x\)), but it is not influential — it sits close enough to where the line would go without it that removing it barely changes the fit. (c) The outlier is near the center of the data in the \(x\)-direction, so it has low leverage. It is neither high-leverage nor influential; its effect on the slope is minimal.
  5. (a) There is a negative, moderate-to-strong, roughly linear relationship between percent of families who own their home and percent of the population living in urban areas. One clear outlier (the state where 100% of the population is urban) sits far to the lower right; the variability in home-ownership also grows as we move from left to right. (b) The outlier at \((100\%, \text{low})\) is horizontally far from the center of the other points, so it has high leverage. It is also influential: excluding it would noticeably change the slope of the regression line.
  6. (a) Since \(R^2 \approx 52\%\), the correlation is \(r = \pm\sqrt{0.52} \approx \pm 0.72\). The sign matches the direction of the association in the scatterplot from an earlier exercise on the same data: taller people tend to be heavier, so \(r\) is positive, \(r \approx +0.72\). (b) The residuals show slightly larger spread above zero than below, so they are not perfectly symmetric — an indication that the normality condition is stressed, though not badly violated. A simple least-squares fit is still reasonable here.
  7. (a) \(r = \pm\sqrt{R^2}\); the scatterplot shows a positive association between poverty rate and murder rate, so \(r\) is positive. Plug in the reported \(R^2\) (roughly 71%) to get \(r \approx +\sqrt{0.71} \approx +0.84\). (b) The residual plot shows no obvious curvature, no fan shape, and no extreme outliers, so Linearity, Normality, and Equal variance all look adequate. Independence is plausible if the 20 metropolitan areas were sampled independently. A simple least-squares fit is appropriate.
  8. (a) The residual spread grows with the predicted value — small residuals for small predicted values, larger residuals for larger predicted values. This is a fan-shaped pattern, and it flags a violation of the Equal-variance (constant variability) condition. (b) Unequal variability does not affect the fit of the line itself — the regression line still models the average heart weight at each body weight. But it does affect the inference (the p-value and CI for the slope), because the pooled SE used in the \(t\)-test underestimates uncertainty in the wide-spread region and overstates it in the narrow region. The reported p-value can be trusted only loosely; if inference is central, a bootstrap CI for the slope is a safer read.

StatLens Exercises

  1. (a) Three features of a good residual plot: (i) random scatter around 0 with no visible trend, (ii) roughly constant spread across \(x\) (or across predicted value), (iii) no outliers sitting far from the rest of the cloud. (b) Three bad patterns: curved pattern → the linear model is mis-specified; the true relationship is non-linear. Fan shape → the equal-variance assumption fails (heteroscedasticity). One huge residual → a possible outlier or data-quality issue worth investigating. (c) Small residuals mean the points sit close to the line, but say nothing about whether the line is the right shape of curve. A mis-specified model (e.g., a line fit to a curved relationship) can have small residuals near the middle of the data range and still be badly wrong. (d) The equal-variance (homoscedasticity) condition has failed. The reported \(t\)-statistic and CI use a single pooled SE, which under-states uncertainty where spread is wide and over-states it where spread is narrow. The \(t\)-CI’s actual coverage is below 95% — a bootstrap CI for the slope is more honest in this situation.
  2. (a) It is not just unreliable, it is conceptually wrong: month 25 lies outside the data range (1–12), and there is no evidence that whatever trend held inside the data continues outside it. Worse, month is periodic (month 13 wraps back to January), so a straight line is the wrong shape of model for the phenomenon in the first place. (b) Trusted: 5 hours (interior to the data). Not trusted: 15 hours (extrapolation — possible but unsupported). Nonsensical: \(-2\) hours (extrapolation and physically impossible). (c) Extrapolating tells you what the fitted line would predict outside the range of \(x\), but it tells you nothing about whether the linear relationship still holds there. The model has no evidence for or against any behavior outside the observed data. (d) The model was fit on homes in the 500–3500 sqft range. The relationship between price and area at very large sizes is almost certainly different (different buyer pool, fewer comparables, non-linear scaling). The prediction interval’s width is computed as if the linear model holds at \(x = 50000\) — but that assumption is itself the dominant source of error, and the reported width does not reflect it.
  3. (a) Leverage measures how far an observation’s \(x\)-value is from \(\bar{x}\) relative to the spread of \(x\). A point at extreme \(x\) sits at the end of the “lever arm” and has the potential to pull the slope a lot — leverage is a position property, decided by \(x\) alone, before \(y\) is even considered. (b) The slope changes very little. A high-leverage point that sits on the fitted line simply confirms the trend; removing it barely moves the fit. High leverage in isolation is not a problem. (c) The slope changes a lot. A point with high leverage and a large residual is called an influential observation: its presence pulls the regression line toward itself, distorting the fit. (d) Before deciding to keep or drop a high-leverage point, ask: (i) Is it a data-quality error? (typo, mis-recorded value, unit mismatch) — if so, correct or remove it and document the reason. (ii) Is it a real but extreme observation from the same population? — if so, report the analysis both with and without the point, and be transparent about how sensitive the conclusion is to its inclusion.

E.22 Chapter 22 — ANOVA for Regression

Solutions to the odd-numbered exercises in Chapter 22 — ANOVA for Regression.

  1. (a) \(df_{\text{Reg}} = 1\), \(df_{\text{Res}} = n - 2 = 58\), \(df_{\text{Total}} = n - 1 = 59\). \(MSR = 442/1 = 442\). \(MSE = 128/58 \approx 2.21\). \(F = MSR/MSE \approx 200.3\). \(SST = SSR + SSE = 442 + 128 = 570\). (b) \(R^2 = SSR/SST = 442/570 \approx 0.775\). About 77.5% of the variability in highway mpg is explained by the linear regression on curb weight. (c) \(H_0: \beta_1 = 0\) vs. \(H_A: \beta_1 \neq 0\). An \(F\)-statistic near 200 on \((1, 58)\) df is enormous compared with the “\(F \approx 1\)” null expectation, so the p-value is effectively 0 — strong evidence that curb weight is linearly associated with highway mpg.
  2. (a) \(t^2 = 3.1^2 = 9.61 = F\) (to two decimal places), confirming the identity numerically. (b) In simple linear regression the model-level \(F\)-test and the slope \(t\)-test are the same test written in different notation: \(F_{1,\,df} = t^2_{\,df}\), so they must always produce identical p-values.
  3. (a) \(F = 0.34\) is below 1, and the null-expected value is near 1. That puts the observed \(F\) in the bulk of the null distribution, so the p-value is large — roughly 0.5 or bigger (a formal computation gives \(p \approx 0.56\)). (b) No evidence in these data of a linear relationship between voltage and bulb lifetime. (c) A non-rejecting \(F\)-test does not prove there is no relationship. Two alternative explanations: (i) the relationship exists but is nonlinear (a curve or threshold effect), which the linear \(F\)-test cannot detect; (ii) a real linear effect exists but the sample of \(n = 40\) has too little power to reach significance — a larger study might discern it.

StatLens Exercises

  1. (a) Both tests state \(H_0: \beta_1 = 0\) vs. \(H_A: \beta_1 \neq 0\). (b) They are the same hypothesis because in simple linear regression there is only one slope to test. In multiple regression the \(F\)-test tests all slopes jointly (\(H_0: \beta_1 = \beta_2 = \cdots = 0\)) while each \(t\)-test tests a single slope at a time; there the two tests diverge. (c) \(t = \pm\sqrt{F} = \pm\sqrt{12.25} = \pm 3.5\). Because \(F_{1,\,48} = t^2_{\,48}\), this \(t\) produces exactly the same p-value.
  2. (a) “About 72% of the variability in the response is explained by the linear model on this predictor.” (b) “72% accurate” is a classification idea, not a regression one; individual predictions are not each 72% correct. The correct statement is that 72% of the variability in \(y\) is explained by the fitted line. (c) \(R^2 = 0.72\) does not tell you: (i) whether the LINE conditions hold — a curved relationship can still produce a high \(R^2\); (ii) whether the model will generalize to new data — \(R^2\) is computed on the fit sample, not held-out data; (iii) whether the fit is practically useful — \(R^2 = 0.72\) might be excellent for one problem and poor for another, depending on how much precision the decision at hand requires.
  3. (a) This is the classic multicollinearity signature: the predictors are so correlated with each other that no single one is uniquely responsible for the fit, but together they explain a discernible share of the variability. The \(F\)-test picks up the joint contribution; the individual \(t\)-tests each ask “does this predictor add something beyond the others?” — and when the predictors are redundant, none does individually. (b) Neither the joint fit nor any single predictor’s unique contribution is discernible from the data. Either there is no real relationship, or the sample is too small for the effects that are present to reach significance — more data would help. (c) With correlated predictors and small samples this pattern can arise, but the model-level \(F\)-test is the more trustworthy summary for whether the model works as a whole; a single significant \(t\) against a non-significant \(F\) is often a false positive on the individual test. Refit with fewer predictors or collect more data before drawing conclusions.

E.23 Chapter 23 — Multiple Regression

Solutions to the odd-numbered exercises in Chapter 23 — Multiple Regression.

  1. (a) \(\widehat{\texttt{interest\_rate}} = 4.31 + 0.041(15) + 0.16(36) + 0.25(2) = 4.31 + 0.615 + 5.76 + 0.50 = 11.19\) (percent). (b) Among borrowers with the same debt-to-income and same loan term, each additional credit check in the last 12 months is associated with an increase of about \(0.25\) percentage points in the predicted interest rate, holding the other predictors constant. (c) In symbols, \(H_0: \beta_{\text{credit\_checks}} = 0\), given that debt_to_income and term are already in the model. In words: after accounting for debt-to-income and loan term, the number of recent credit checks has no linear relationship with interest rate. A small p-value means credit checks add explanatory power beyond what those two predictors already provide.
  2. (a) Two predictors: Years of experience (VIF \(= 8.7\)) and Years at company (VIF \(= 9.1\)). Both are close to the rule-of-thumb threshold of 10, which flags severe multicollinearity. In context, someone with more overall experience tends also to have more time at the current company — the two variables carry overlapping information. (b) When two predictors move together, the model cannot separate their individual contributions cleanly; the standard error on each coefficient inflates, so the estimated slope can shift wildly with small changes in the data. That makes “the effect of years of experience, holding years at company constant” hard to pin down. (c) Drop one of the two collinear predictors, OR combine them into a single derived variable (for example, total experience or a ratio like years-at-company / years-of-experience). A third option is to keep both but stop interpreting the individual slopes and use the model only for prediction.
  3. (a) No. A p-value of \(0.28\) is not evidence that the predictor helps; combined with the strong pairwise association with an existing predictor, adding it risks multicollinearity without gaining useful signal. (b) Yes — subject-matter knowledge is a legitimate reason to include a predictor. Statistical non-significance is not evidence of “no effect”; it just means the sample did not resolve one. If theory says the variable matters, keeping it in yields an honest estimate of its coefficient (and its uncertainty) given the other predictors. (c) The residual standard error measures how much variability in the response remains unexplained. Essentially unchanged \(s\) means the predictor is not reducing residual variability — it is not earning its degree of freedom in a predictive sense.

StatLens Exercises

  1. (a) \(\texttt{interest\_rate} = \beta_0 + \beta_1 \cdot \texttt{debt\_to\_income} + \beta_2 \cdot \texttt{term} + \beta_3 \cdot \texttt{credit\_checks} + \varepsilon\). (b) \(H_0\): none of the three predictors is linearly related to interest rate (\(\beta_1 = \beta_2 = \beta_3 = 0\)). \(H_A\): at least one predictor is linearly related to interest rate. (c) \(H_0: \beta_1 = 0\), given term and credit_checks are already in the model. This asks a conditional question — does debt_to_income add anything above and beyond what term and credit checks already explain? The \(F\)-test asks the unconditional joint question — does any predictor contribute at all?
  2. (a) StatLens reports the fitted equation directly. It will read \(\widehat{\text{head\_l}} = b_0 + b_1 \cdot \text{total\_l} + b_2 \cdot \text{skull\_width}\); students should copy the actual numeric coefficients from the tool. (b) Each slope is the mean change in head length per one-unit change in that predictor, holding the other predictor fixed. Report the actual numbers and units (mm per mm) from the tool. (c) The multi-predictor total-length slope is smaller in magnitude than the simple-regression slope. Larger possums tend to have both longer total lengths and wider skulls, so the simple slope absorbed some of the skull-width effect. The multi-predictor slope isolates the total-length effect after skull width is accounted for. (d) The model-level \(F\)-statistic is large and the p-value is essentially zero — both predictors carry real signal about head length.
  3. (a) The \(F\)-statistic is large enough that its p-value falls below \(0.05\), while every individual \(t\)-test has a p-value well above \(0.05\). This pattern is the fingerprint of multicollinearity: the joint contribution of the predictors is real, but no single predictor is uniquely responsible for it. (b) Yes — the pairwise scatterplots show one or more strongly linear pairs. In expert mode, at least two VIFs will be well above 5, confirming what the scatterplots suggested. (c) The model is trustworthy for prediction: put in the predictor values, get an interest-rate estimate. It is not trustworthy for interpretation: statements like “predictor \(x_1\) specifically drives the response” are unreliable because \(x_1\) can be swapped for a linear combination of the other predictors without hurting the fit. (d) Rejecting the model based on individual \(t\)-tests would throw out a jointly informative model just because its predictors overlap. Better responses are to combine the collinear predictors, drop one of them, or accept the model for prediction and refuse to interpret individual slopes.

E.24 Chapter 24 — Probability Rules

Solutions to the odd-numbered exercises in Chapter 24 — Probability Rules.

  1. (a) False. This is the gambler’s fallacy: successive free throws are essentially independent, so her next attempt still has about a 70% chance of going in — past misses do not “compensate.” (b) False. Rolling a 4 is a subset of rolling an even number ({2, 4, 6}); the two events occur together whenever she rolls a 4, so they are not mutually exclusive. (c) True. “It rains” and “it does not rain” are complementary events, and complementary probabilities sum to 1.
  2. In every version, the sample proportion of heads is more concentrated near 0.50 with \(n = 100\) than with \(n = 10\) (law of large numbers). Choose the flip count that pushes the proportion into the winning region. (a) Larger than 0.60 — 10 flips (needs a lucky excursion above 0.60; unlikely at \(n = 100\)). (b) Larger than 0.40 — 100 flips (with \(n = 100\) the proportion is almost always above 0.40). (c) Between 0.40 and 0.60 — 100 flips (concentration inside the middle interval). (d) Smaller than 0.30 — 10 flips (again needs a lucky excursion, this time to the low side).
  3. (a) \(P(\text{all tails}) = (1/2)^{10} = 1/1024 \approx 0.001\). (b) \(P(\text{all heads}) = (1/2)^{10} = 1/1024 \approx 0.001\). (c) \(P(\text{at least one tails}) = 1 - P(\text{all heads}) = 1 - 1/1024 = 1023/1024 \approx 0.999\).
  4. Let \(I =\) Independent, \(S =\) swing voter. \(P(I) = 0.35\), \(P(S) = 0.23\), \(P(I \text{ and } S) = 0.11\). (a) Not disjoint — 11% of voters are both. (b) Venn: only \(I\) = \(0.35 - 0.11 = 0.24\); both = \(0.11\); only \(S\) = \(0.23 - 0.11 = 0.12\); neither = \(1 - 0.24 - 0.11 - 0.12 = 0.53\). (c) \(P(I \text{ only}) = 0.24\), or 24%. (d) \(P(I \text{ or } S) = 0.35 + 0.23 - 0.11 = 0.47\), or 47%. (e) \(P(\text{neither}) = 1 - 0.47 = 0.53\), or 53%. (f) \(P(S \mid I) = 0.11/0.35 \approx 0.314 \ne P(S) = 0.23\), so being a swing voter is not independent of being an Independent.
  5. (a) Independent. Two randomly selected students’ outcomes are unrelated (assuming independent grading), so one earning an A does not change the probability the other does. Not disjoint — both can earn A’s. (b) Neither. Study partners’ grades are positively related (shared study time, materials), so their outcomes are not independent. They also can both earn A’s, so not disjoint. (c) No. Being able to occur at the same time only rules out disjoint; independence is a separate property (about whether one event’s occurrence changes the probability of the other).
  6. (a) Men with at least a Bachelor’s: \(0.16 + 0.09 = 0.25\). (b) Women with at least a Bachelor’s: \(0.17 + 0.09 = 0.26\). (c) Assuming the husband’s and wife’s education levels are independent, \(P(\text{both} \ge \text{Bachelor's}) = 0.25 \times 0.26 = 0.065\). (d) The independence assumption is not reasonable — people tend to partner with others of similar educational attainment (assortative mating), so the true probability is likely higher than 0.065.
  7. With \(P(A) = 0.3\) and \(P(B) = 0.7\): (a) No. Without knowing independence or the joint probability, \(P(A \text{ and } B)\) is not determined by \(P(A)\) and \(P(B)\) alone. (b) Assuming independence: (c) \(P(A \text{ and } B) = 0.3 \times 0.7 = 0.21\). (d) \(P(A \text{ or } B) = 0.3 + 0.7 - 0.21 = 0.79\). (e) \(P(A \mid B) = P(A) = 0.3\) (independence means conditioning on \(B\) does not change \(A\)’s probability). (f) If \(P(A \text{ and } B) = 0.1\): independence would require \(P(A) \times P(B) = 0.21 \ne 0.1\), so \(A\) and \(B\) are not independent. (g) \(P(A \mid B) = P(A \text{ and } B) / P(B) = 0.1 / 0.7 \approx 0.143\).
  8. Let \(W =\) believes earth is warming, \(L =\) liberal Democrat. From the joint table, \(P(W) = 0.60\), \(P(L) = 0.20\), \(P(W \text{ and } L) = 0.18\). (a) Not mutually exclusive — 18% are both. (b) \(P(W \text{ or } L) = 0.60 + 0.20 - 0.18 = 0.62\). (c) \(P(W \mid L) = 0.18/0.20 = 0.90\). (d) For conservative Republicans (\(P(W \text{ and } CR) = 0.11\), \(P(CR) = 0.33\)): \(P(W \mid CR) = 0.11/0.33 \approx 0.33\). (e) Not independent. \(P(W \mid L) = 0.90\) while \(P(W \mid CR) \approx 0.33\) and \(P(W) = 0.60\) — the belief probability changes sharply with party/ideology. (f) \(P(\text{Mod/Lib Rep} \mid \text{not } W) = 0.06 / 0.34 \approx 0.18\).
  9. Column totals from the table: 248 males, 252 females (500 total). (a) Not mutually exclusive — six women pick Five Guys. (b) \(P(\text{In-N-Out} \mid \text{male}) = 162/248 \approx 0.653\). (c) \(P(\text{In-N-Out} \mid \text{female}) = 181/252 \approx 0.718\). (d) Assuming the man’s and woman’s preferences are independent, \(P(\text{both In-N-Out}) = 0.653 \times 0.718 \approx 0.469\). The independence assumption is questionable — dating partners often share tastes (or influence each other), so preferences may be correlated. (e) \(P(\text{Umami or female}) = P(U) + P(F) - P(U \text{ and } F) = 6/500 + 252/500 - 1/500 = 257/500 = 0.514\).
  10. Let \(T =\) attends tutoring, \(P =\) passes. \(P(T) = 0.70\), \(P(P \mid T) = 0.90\), \(P(P \mid T^c) = 0.72\). (a) \(P(T \text{ and } P) = P(P \mid T) P(T) = 0.90 \times 0.70 = 0.63\). (b) By the law of total probability, \(P(P) = P(P \mid T) P(T) + P(P \mid T^c) P(T^c) = 0.90(0.70) + 0.72(0.30) = 0.63 + 0.216 = 0.846\). (c) By Bayes’ rule, \(P(T \mid P) = P(T \text{ and } P) / P(P) = 0.63/0.846 \approx 0.745\).
  11. Let \(D =\) has lupus, and \(+ =\) tests positive. \(P(D) = 0.02\), \(P(+ \mid D) = 0.98\), and “74% accurate if the person does not have the disease” means \(P(- \mid D^c) = 0.74\), so \(P(+ \mid D^c) = 0.26\). Total probability: \(P(+) = 0.02(0.98) + 0.98(0.26) = 0.0196 + 0.2548 = 0.2744\). Bayes’ rule: \(P(D \mid +) = 0.0196/0.2744 \approx 0.071\). Even after a positive test, there is only about a 7% chance the patient has lupus — the disease is so rare that false positives from the (much larger) healthy population overwhelm the true positives. In that sense, “it’s never lupus” is a defensible first guess.
  12. With replacement, each draw sees the original 10-marble bag (5 red, 3 blue, 2 orange). (a) \(P(\text{first blue}) = 3/10 = 0.30\). (b) With replacement, the second draw is independent of the first: \(P(\text{second blue}) = 3/10 = 0.30\). (c) Same reasoning — with replacement, prior draws don’t change the composition: \(P(\text{second blue}) = 3/10 = 0.30\). (d) \(P(\text{two blue}) = (3/10)^2 = 9/100 = 0.09\). (e) Yes, the draws are independent: replacement resets the bag, so what you drew before does not change the probability of what you draw next.
  13. Without replacement from the 10-chip bag (5 red, 3 blue, 2 orange). (a) After removing 1 blue, 9 chips remain with 2 blue: \(P(\text{second blue}) = 2/9\). (b) After removing 1 orange, 9 chips remain with 3 blue: \(P(\text{second blue}) = 3/9 = 1/3\). (c) \(P(\text{two blue}) = (3/10)(2/9) = 6/90 = 1/15 \approx 0.067\). (d) No — the composition of the bag changes after each draw, so the probability of a color on the second draw depends on what was drawn first.
  14. Total students: 24 with 7 jeans, 4 shorts, 8 skirts, and \(24 - 7 - 4 - 8 = 5\) leggings. Draw 3 without replacement; want exactly 1 leggings and 2 jeans: \[P = \frac{\binom{5}{1}\binom{7}{2}}{\binom{24}{3}} = \frac{5 \times 21}{2024} = \frac{105}{2024} \approx 0.052.\] Equivalently, \(3 \times (5/24)(7/23)(6/22) \approx 0.052\) (three orderings for which slot holds the leggings).
  15. Using the joint counts (Yes-coverage row from the table plus the No-coverage row from the source: Excellent 459, Very Good 524, Good 596, Fair 262, Poor 132, No-total 1,973). Grand total \(= 17{,}476 + 1{,}973 = 19{,}449\). (a) \(P(\text{excellent AND no coverage}) = 459/19{,}449 \approx 0.024\). (b) \(P(\text{excellent OR no coverage}) = P(\text{excellent}) + P(\text{no coverage}) - P(\text{both}) = (4{,}198 + 459)/19{,}449 + 1{,}973/19{,}449 - 459/19{,}449 = 6{,}171/19{,}449 \approx 0.317\).
  16. Let \(M =\) manufactured on Machine A, and \(D_3 =\) all three inspected parts are defective. Prior \(P(A) = 0.40\), \(P(B) = 0.60\); defect rates \(P(\text{defective} \mid A) = 0.02\) and \(P(\text{defective} \mid B) = 0.05\). Assuming independence of the three parts from the same machine, \(P(D_3 \mid A) = 0.02^3 = 8 \times 10^{-6}\) and \(P(D_3 \mid B) = 0.05^3 = 1.25 \times 10^{-4}\). (a) By Bayes’ rule, \[P(A \mid D_3) = \frac{P(D_3 \mid A) P(A)}{P(D_3 \mid A) P(A) + P(D_3 \mid B) P(B)} = \frac{0.4 \times 8 \times 10^{-6}}{0.4 \times 8 \times 10^{-6} + 0.6 \times 1.25 \times 10^{-4}} \approx 0.041.\] About 4.1%. (b) Three consecutive defects are much more likely under Machine B (defect rate 5%) than under Machine A (defect rate 2%): the likelihood ratio is about \(0.05^3 / 0.02^3 = 15.6\), which sharply outweighs the 40/60 prior in Machine B’s favor.

E.25 Chapter 25 — Random Variables

Solutions to the odd-numbered exercises in Chapter 25 — Random Variables.

  1. (a) If we view $X = $ number of smokers in a random sample of 100 students as \(\text{Binomial}(n = 100, p = 0.13)\), then \(E(X) = np = 100 \times 0.13 = 13\) smokers. (b) No. The 27 students waiting outside the gym at 8:55 am on a Saturday are not a random sample of the university — they self-selected into an early-morning exercise activity, so the smoking rate in this group is likely lower than the campus-wide 13%. The binomial expected-value formula requires the trials to be identically distributed random draws, which fails here.
  2. Let \(W\) be the payout. With \(\binom{12}{3} = 220\) equally likely three-marble subsets, \(P(\text{3 red}) = \binom{5}{3}/220 = 10/220\), \(P(\text{3 blue}) = \binom{4}{3}/220 = 4/220\), and \(P(W = 0) = 206/220\).

    \(a\) $E(W) = 40 \cdot \frac{10}{220} + 25 \cdot \frac{4}{220} + 0 \cdot \frac{206}{220} = \frac{500}{220} \approx \$2.27$.
    $E(W^2) = 40^2 \cdot \frac{10}{220} + 25^2 \cdot \frac{4}{220} = \frac{18{,}500}{220} \approx 84.09$, so
    $\text{Var}(W) \approx 84.09 - (2.27)^2 \approx 78.93$ and $\text{SD}(W) \approx \$8.88$.
    
    (b) $N = W - 5$ shifts the mean by $-5$ and leaves the spread unchanged: $E(N) \approx -\$2.73$ and $\text{SD}(N) \approx \$8.88$.
    
    (c) No. The expected net profit is about $-\$2.73$ per play, so in the long run a risk-neutral player loses roughly \$2.73 each time.</li>
  3. The eight bins have counts \(29, 32, 21, 25, 12, 15, 5, 4\), totalling \(n = 143\) cats.

    \(a\) The two bins below $2.5$ kg contain $29 + 32 = 61$ cats, so about $61/143 \approx 0.43$ (43%).
    
    (b) The bin from $2.5$ to $2.75$ kg contains $21$ cats, so about $21/143 \approx 0.15$ (15%).
    
    (c) The three bins from $2.75$ to $3.5$ kg contain $25 + 12 + 15 = 52$ cats, so about $52/143 \approx 0.36$ (36%).</li>
  4. A valid probability distribution requires (i) every probability in \([0, 1]\) and (ii) probabilities summing to \(1\).

    (b) $0, 0, 1, 0, 0$: sums to $1$ and all values are in $[0, 1]$ --- **valid** (a degenerate distribution that assigns all probability to one grade).
    
    (c) $0.3, 0.3, 0.3, 0, 0$: sums to $0.9 \neq 1$ --- **invalid**.
    
    (d) $0.3, 0.5, 0.2, 0.1, -0.1$: sums to $1$, but $-0.1$ is not a legal probability --- **invalid**.
    
    (e) $0.2, 0.4, 0.2, 0.1, 0.1$: sums to $1$ and all values are in $[0, 1]$ --- **valid**.
    
    (f) $0, -0.1, 1.1, 0, 0$: sums to $1$, but $-0.1 < 0$ and $1.1 > 1$ --- **invalid**.</li>

:::

E.26 Chapter 26 — Expected Value and Variance

Solutions to the odd-numbered exercises in Chapter 26 — Expected Value and Variance.

  1. Each scenario is equally likely, so \(P = 1/3\) for each. The expected return is \(E(R) = \tfrac{1}{3}(18) + \tfrac{1}{3}(9) + \tfrac{1}{3}(-12) = \tfrac{15}{3} = 5\%\).
  2. Let \(W\) = winnings on a $1 bet on red. \(W = \$1\) with probability \(18/38 = 9/19\) and \(W = -\$1\) with probability \(20/38 = 10/19\). \(E(W) = 1 \cdot \tfrac{18}{38} + (-1) \cdot \tfrac{20}{38} = -\tfrac{2}{38} = -\tfrac{1}{19} \approx -\$0.053\). For the standard deviation, \(E(W^2) = (1)^2 \cdot \tfrac{18}{38} + (-1)^2 \cdot \tfrac{20}{38} = 1\), so \(\text{Var}(W) = 1 - (-1/19)^2 = 1 - 1/361 = 360/361\) and \(\text{SD}(W) = \sqrt{360/361} \approx \$1.00\). On average you lose about 5 cents per $1 wagered, with a typical bet-to-bet swing of about a dollar.
  3. Let \(C\) = price of coffee and \(M\) = price of muffin, with \(\mu_C = \$1.40\), \(\sigma_C = \$0.30\), \(\mu_M = \$2.50\), \(\sigma_M = \$0.15\), and \(C \perp M\). (a) Daily total \(D = C + M\). \(E(D) = 1.40 + 2.50 = \$3.90\); \(\text{Var}(D) = (0.30)^2 + (0.15)^2 = 0.09 + 0.0225 = 0.1125\); \(\text{SD}(D) = \sqrt{0.1125} \approx \$0.34\). (b) Weekly total \(W = D_1 + D_2 + \cdots + D_7\) (days independent). \(E(W) = 7(3.90) = \$27.30\); \(\text{Var}(W) = 7(0.1125) = 0.7875\); \(\text{SD}(W) = \sqrt{0.7875} \approx \$0.89\). Note that the mean scales by 7 but the SD scales by only \(\sqrt{7} \approx 2.65\) — the reason averaging (or, equivalently, larger \(n\)) shrinks standard error.
  4. (a) By independence, \(\text{Var}(X_1 + X_2) = \text{Var}(X_1) + \text{Var}(X_2) = 25 + 25 = 50\). Since \(\bar{X} = (X_1 + X_2)/2\), \(\text{Var}(\bar{X}) = \tfrac{1}{4}\text{Var}(X_1 + X_2) = \tfrac{1}{4}(50) = 12.5\). (b) For \(n = 2\), \(\text{Var}(\bar{X}) = \sigma^2/2\).
  5. (a) In general, \(\text{Var}(\bar{X}) = \sigma^2/n\) (the sample-mean variance formula). This follows from \(\bar{X} = (1/n)\sum X_i\): \(\text{Var}(\bar{X}) = (1/n)^2 \sum \text{Var}(X_i) = (1/n^2)(n\sigma^2) = \sigma^2/n\), using independence. (b) With \(\sigma^2 = 36\) and \(n = 16\), \(\text{Var}(\bar{X}) = 36/16 = 2.25\) and \(\text{SD}(\bar{X}) = \sqrt{2.25} = 1.5\).

E.27 Chapter 27 — Binomial Distribution

Solutions to the odd-numbered exercises in Chapter 27 — Binomial Distribution.

  1. (a) Yes. The four binomial conditions are met: (i) \(n = 10\) trials, (ii) each 18–20 year old either did or did not consume alcohol, (iii) the success probability \(p = 0.697\) is constant across a random sample (10 is a tiny fraction of the U.S. 18–20 population, so the 10% rule for independence holds), (iv) trials are independent. (b) \(P(X = 6) = \binom{10}{6}(0.697)^6(0.303)^4 \approx 210 \cdot 0.1147 \cdot 0.00843 \approx 0.203\). (c) “Exactly 4 have not consumed” is the same event as “exactly 6 have consumed,” so the answer is again \(\approx 0.203\). (d) With \(n = 5\), \(P(X \le 2) = P(0) + P(1) + P(2) \approx 0.0026 + 0.0294 + 0.135 \approx 0.167\). (e) \(P(X \ge 1) = 1 - P(X = 0) = 1 - (0.303)^5 \approx 1 - 0.0026 \approx 0.997\).
  2. (a) \(\mu = np = 50(0.70) = 35\) and \(\sigma = \sqrt{np(1-p)} = \sqrt{50(0.70)(0.30)} = \sqrt{10.5} \approx 3.24\). (b) Yes. \(45\) is \((45 - 35)/3.24 \approx 3.1\) standard deviations above the mean — well into the tail. (c) Using the normal approximation with continuity correction, \(z = (44.5 - 35)/3.24 \approx 2.93\), so \(P(X \ge 45) \approx P(Z \ge 2.93) \approx 0.0017\). The tiny probability confirms the intuition in (b) — seeing \(45\) or more would be surprising.
  3. With \(n = 3\) spins and \(p = 1/4\) for the color of interest: (a) \(P(\text{at least one red}) = 1 - (3/4)^3 = 1 - 27/64 = 37/64 \approx 0.578\). (b) \(P(\text{exactly 2 reds}) = \binom{3}{2}(1/4)^2(3/4) = 9/64 \approx 0.141\). (c) \(P(\text{exactly 1 blue}) = \binom{3}{1}(1/4)(3/4)^2 = 27/64 \approx 0.422\). (d) \(P(\text{at most 2 greens}) = 1 - P(\text{3 greens}) = 1 - (1/4)^3 = 63/64 \approx 0.984\).
  4. Let \(p = 0.125\) be the probability of green eyes. (a) \(P(\text{1st green, 2nd not}) = 0.125 \cdot 0.875 \approx 0.109\). (b) \(P(\text{exactly one of two green}) = \binom{2}{1}(0.125)(0.875) \approx 0.219\). (c) \(P(\text{exactly 2 of 6 green}) = \binom{6}{2}(0.125)^2(0.875)^4 \approx 15 \cdot 0.0156 \cdot 0.586 \approx 0.137\). (d) \(P(\text{at least one of 6 green}) = 1 - (0.875)^6 \approx 1 - 0.449 \approx 0.551\). (e) \(P(\text{1st green is the 4th child}) = (0.875)^3(0.125) \approx 0.670 \cdot 0.125 \approx 0.0837\). (f) With \(p_{\text{brown}} = 0.75\), \(\mu = 6(0.75) = 4.5\) and \(\sigma = \sqrt{6(0.75)(0.25)} \approx 1.06\). Only 2 brown-eyed children is roughly \(2.4\) SD below the mean; \(P(X \le 2) \approx 0.038\). Yes, that would be considered unusual.
  5. (a) With 8 distinct songs and no repeats, there are \(8! = 40{,}320\) orderings. (b) By symmetry, each song is equally likely to be first, so \(P(\text{Track 1 first}) = 1/8 = 0.125\). (c) Fix Track 1 in slot 1 and Track 2 in slot 8; the remaining 6 songs can fill slots 2–7 in \(6!\) ways. So \(P = 6!/8! = 1/(8 \cdot 7) = 1/56 \approx 0.0179\).
  6. (a) \(\mu = np = 54(0.93) = 50.22\). (b) \(\sigma = \sqrt{54(0.93)(0.07)} = \sqrt{3.5154} \approx 1.875\). (c) With continuity correction, \(P(X > 50) = P(X \ge 51) \approx P\!\left(Z \ge \dfrac{50.5 - 50.22}{1.875}\right) = P(Z \ge 0.149) \approx 0.441\). (d) Even though the airline sold only 54 tickets under the assumption that about 7% will not show, there is roughly a \(44\%\) chance the flight is overbooked — expect bumped passengers on roughly two flights out of every five. Selling 54 tickets is quite aggressive.

StatLens Exercises

  1. (a) Yes — all four conditions met with \(n = 50\) and \(p = 0.5\). (b) No — drawing without replacement changes \(p\) after each draw, violating the constant-\(p\) and independence conditions. This is the hypergeometric setting; when the sample is small relative to the deck, the binomial is a reasonable approximation. (c) Yes (approximately)\(n = 100\) fixed, two outcomes, and \(p\) is essentially constant because 100 is a small fraction of the adult population (the 10% rule for independence). (d) No\(n\) is not fixed; the number of trials is itself random. This is the geometric distribution. (e) Yes (approximately) — the four conditions hold as long as bulbs are independent (no shared manufacturing defect that would make failures cluster).
  2. (a) At \(n = 20\), \(p = 0.05\) the PMF is strongly right-skewed; most mass sits at \(0\)\(3\) successes. The normal overlay is a poor fit — it extends to negative values and is symmetric, while the binomial cannot go below \(0\) and is asymmetric. (b) At \(n = 200\), \(p = 0.05\) the PMF is approximately bell-shaped and centered near \(10\); the normal overlay fits well. (c) Case (a): \(np = 1\) fails the rule. Case (b): \(np = 10\) and \(n(1-p) = 190\) — both pass. (d) The rule needs both pieces because a distribution can be skewed either way. At \(p = 0.99\), \(n = 100\): \(np = 99\) passes easily, but \(n(1-p) = 1\) fails — the distribution is squished against its upper bound and is left-skewed. Checking only \(np\) would miss this.
  3. (a) Expected wins \(= np = 45\). Probability of coming out ahead: \(P(X \ge 51) = 1 - P(X \le 50) \approx 0.134\), about \(13\%\). (b) The law of large numbers says your average per-round result will converge to the expected value: \(E[X] - n/2 = 45 - 50 = -\$5\) per round (equivalently \(-\$0.05\) per trial). Random “good” rounds are cancelled out by random “bad” rounds in the long run. (c) The expected loss is small per round, but with independent rounds the LLN guarantees that total losses grow linearly with the number of rounds — exactly the opposite of what the player wants to believe. The “small per round” framing hides the fact that the house edge is paid every round. (d) The probability of coming out ahead shrinks as \(n\) grows: about \(13\%\) at \(n = 100\), well under \(1\%\) at \(n = 1000\), essentially \(0\) at \(n = 10{,}000\). As \(n\) grows, \(\hat{p}\) concentrates around the true \(p = 0.45\), so \(\hat{p} > 0.5\) becomes vanishingly rare.

E.28 Chapter 28 — More About the Normal Distribution

Solutions to the odd-numbered exercises in Chapter 28 — More About the Normal Distribution.

  1. (a) Approximately normal — a straight-line QQ plot is exactly what we hope to see; the observed quantiles match the normal quantiles at every point. (b) Right-skewed — the upper tail extends farther than a normal would predict, so the largest observed values exceed the normal quantiles and the points rise above the line. (c) Left-skewed — the lower tail extends farther than normal, so the smallest observed values fall below the normal quantiles and the points drop below the line on the left. (d) Heavy-tailed — both extremes are more extreme than a normal would predict (typical of a \(t\)-distribution with low df, or data with occasional outliers on both sides).
  2. (a) Consistent with normality. With \(n = 15\) we expect some scatter around the line even when the data truly are normal — random sampling variability alone produces wobble at small \(n\). The signal that matters is the presence or absence of a systematic curve (S-shape, monotonic bend at one end), not the amount of wobble. (b) Yes, the conclusion changes. With \(n = 500\) random sampling variability is much smaller and the plot should look much cleaner; persistent scatter at that sample size would indicate a real departure from normality (mixture distribution, heteroscedasticity, etc.). (c) No — the reasoning is unsound. The \(t\)-test is robust to moderate departures from normality, especially with modest-to-large \(n\); perfect linearity of the QQ plot is not required. The right questions are whether the departure is severe (clear skew or heavy tails) and whether \(n\) is small enough for departures to matter. Modest wobble at any \(n\) is not a reason to abandon the \(t\)-test.
  3. (a) Agree. Both diagnostics point to strong right skew. The histogram makes the shape visible directly; the QQ plot’s upward curve at the right end confirms an over-long upper tail relative to a normal. (b) A log transformation (\(\log(x + \varepsilon)\), or \(\sqrt{x}\) for zero-heavy data) is a natural choice for right-skewed positive data — it compresses the long upper tail and expands the crowded region near zero, often producing an approximately symmetric transformed distribution. (c) Cautiously trust the CI, with the caveat that “the mean” may not be the interesting quantity. With \(n = 300\) the CLT makes the sampling distribution of \(\bar{x}\) approximately normal even from skewed data, so the \(t\) CI’s coverage is roughly correct. The deeper caveat: for very skewed data the mean is often not the best summary of a typical day — the median or another percentile may better answer the substantive question.

StatLens Exercises

  1. (a) About 0.683 (68.3%) — the “68%” rule is a rounded version of the exact area. (b) About 0.954 (95.4%) — again, the rule rounds; to get exactly 95% you would need \(\pm 1.96\) SDs, not \(\pm 2\). (c) About 0.997 (99.7%) — essentially exact here. (d) The curve is the truth; the empirical rule is a memorable approximation. When precision matters (a CI at exactly 95%, for instance) use the curve — or the tool — directly, not the \(\pm 2\) shortcut.
  2. (a) Right-tail area = 0.10 → \(z \approx 1.28\) → score \(= 500 + 1.28 \cdot 100 = 628\). (b) Left-tail area = 0.25 → \(z \approx -0.67\) → score \(= 500 - 67 = 433\). (c) 75th percentile \(\approx 500 + 0.67 \cdot 100 = 567\), so \(\text{IQR} = 567 - 433 = 134\) points. (d) Not surprising. The 90th percentile is at 628 (from part a). A 50-point band is half a standard deviation, and near the cutoff the normal density is still high, so a 50-point window can contain many rejected applicants. With \(\sigma = 100\), “near the cutoff” really is a wide window.
  3. (a) Height is roughly the sum of many small genetic and developmental contributions, each acting independently. The Central Limit Theorem for sums of many small independent effects delivers approximate normality. (b) Hourly wages are strongly right-skewed — most workers earn modest hourly rates, with a long upper tail from high earners (surgeons, executives) pulling the mean well above the median. Expect a shape that peaks low and trails off far to the right. (c) A normal model would under-predict the probability of very high wages (the normal’s symmetric tails thin out too quickly) and over-predict wages near the mean (a symmetric peak where the actual distribution is asymmetric). (d) The normal still describes the sampling distribution of \(\bar{x}\) by the CLT, even when the population is skewed — provided \(n\) is large enough. You can build a normal-based CI for the mean wage even though individual wages are not normal. Inference about individual wages (e.g., “what wage is at the 95th percentile?”) requires a different model.

E.29 Chapter 29 — Type II Error and Power

Solutions to the odd-numbered exercises in Chapter 29 — Type II Error and Power.

  1. (a) Scenario (i) is higher. A smaller sample size gives a less precise estimate of \(\hat{p}\), so its standard error is larger. (b) Scenario (i) is higher. Increasing the confidence level widens the interval, which requires a larger margin of error. (c) They are equal. Once the \(Z\)-statistic is fixed, the p-value depends only on the tail area beyond it — the sample size is already baked into the \(Z\)-statistic itself. (d) Scenario (i) is higher. Lowering \(\alpha\) makes \(H_0\) harder to reject, so when \(H_A\) is truly in force we are more likely to fail to detect it (a larger Type II error rate, \(\beta\)).
  2. True. As the sample size grows the standard error shrinks toward zero. With a small enough standard error, even a tiny gap between the observed point estimate and the null value produces a large \(Z\)-statistic and a small p-value. Statistical discernibility is not the same as practical importance — a discernible difference can still be too small to matter.

StatLens Exercises

  1. (a) \(\beta\) is shaded under the alternative curve, in the region below the critical value (the “do not reject \(H_0\)” region). It lives under \(H_A\) because \(\beta\) is defined as the probability of failing to reject when the alternative is true. (b) “\(\beta = 0.14\)” means that if the alternative is really in force at this effect size, this experiment has a 14% chance of missing the effect — a 14% chance of concluding ‘no evidence’ when there truly is an effect. (c) Decreasing \(\delta\) to \(0.5\) makes \(\beta\) rise sharply. A smaller effect pushes the alternative curve closer to the null, so much more of the alternative distribution falls into the do-not-reject region. (d) Decreasing \(\sigma\) to \(1\) makes \(\beta\) fall. Less noise sharpens both curves and pulls their peaks apart, so less of the alternative distribution overlaps the do-not-reject region.
  2. (a) At \(\delta = 0.3,\ n = 30,\ \sigma = 1,\ \alpha = 0.05\) the power is roughly \(0.50\) — essentially a coin flip. A trial with this little power is not informative about whether the effect is real: “no detected effect” is exactly what we’d expect half the time even if the treatment truly worked at \(\delta = 0.3\). (Recall the \(\ge 0.80\) rule of thumb for adequate power.) (b) At \(\delta = 1\) the power is essentially \(1.0\). A non-result at this level of power is informative — it rules out effects that large with high confidence. (c) “No evidence of effect” means the data did not push us to reject \(H_0\); this can happen because there truly is no effect or because the study lacked power to detect it. “Evidence of no effect” is a positive claim that any effect must be small, and requires enough power to have detected a meaningful effect if it existed. The two are not interchangeable. (d) Something like: “We failed to detect an effect of \(\delta \ge 1\) (power \(\ge 0.95\) against effects that large), but were underpowered to detect smaller effects (power \(\approx 0.50\) against \(\delta = 0.3\)). A replication with \(n \ge 200\) would be needed to draw firm conclusions about smaller, clinically meaningful effects.”

E.30 Chapter 30 — Statistical Power

Solutions to the odd-numbered exercises in Chapter 30 — Statistical Power.

  1. (a) (i) is larger. \(SE(\hat{p}) = \sqrt{\hat{p}(1-\hat{p})/n}\) shrinks as \(n\) grows, so the smaller \(n = 125\) gives the larger standard error. (b) (i) is larger. A higher confidence level requires a larger \(z^\star\), which widens the margin of error. (c) Equal. The p-value depends only on the \(Z\)-statistic and the direction of the test — not on the sample size that produced it. (d) (ii) is larger. Increasing the discernibility level (\(\alpha\)) enlarges the rejection region under \(H_0\), which shrinks the rejection region under \(H_A\) and therefore raises \(\beta\).
  2. True. As \(n\) grows, the standard error of the point estimate shrinks like \(1/\sqrt{n}\), so the test statistic grows for a fixed observed difference and the p-value falls. With a large enough sample, a difference that is far too small to matter in context can still be flagged as statistically discernible — which is why practical importance (effect size, context) must be judged separately from a p-value.

StatLens Exercises

  1. (a) The left curve is the null distribution (centered at \(\mu = 0\)); the right curve is the alternative (centered at \(\mu = \delta = 1\)). Each curve is centered at the mean it assumes. (b) \(\alpha\) is the right-tail area of the null curve past the critical value (the false-alarm region under \(H_0\)). \(\beta\) is the area of the alternative curve to the left of the critical value (the missed-detection region under \(H_A\)). Power \(= 1 - \beta\) is the area of the alternative curve to the right of the critical value. (c) At \(\alpha = 0.05\), \(n = 30\), \(\delta = 1\), \(\sigma = 2\), the tool reports power \(\approx 0.756\). (d) “If a real effect of this size exists, this experiment has only a 50/50 chance of detecting it” — roughly a coin flip’s worth of detection ability.
  2. (a) With 500 simulated tests, the empirical reject rate is close to the analytic power of \(\approx 0.76\), typically within a few percentage points of Monte Carlo error. (b) Yes. Monte Carlo error scales like \(1/\sqrt{N}\), so 5000 sims tighten the estimate by roughly a factor of 3 compared to 500. (c) Under a real effect, p-values cluster near 0. Under \(H_0\) they are uniform on \([0, 1]\); under \(H_A\) the distribution shifts toward small values, and the cluster gets denser as the effect size or \(n\) grows. (d) (i) Trust calibration: the analytic formula assumes normality, equal variances, and independence; simulating with the actual generating process tells you whether the analytic power is honest when those assumptions bend. (ii) Intuition building: watching 500 individual studies that “ought to detect a real effect” but about 25% still miss it teaches the meaning of power more viscerally than reading “power \(= 0.76\)” off a tool.
  3. (a) About 80 labs report a statistically discernible result (power \(= 0.8\) means an 80% rejection rate when \(H_A\) is true), and about 20 report a non-result. (b) Publication bias. If only the 80 “discernible” labs publish, the literature is enriched for positive findings, and a reader who sees only those 80 papers overestimates the strength of the effect. Meta-analyses that draw only on published studies inherit the same bias. (c) A meta-analysis of all 100 labs sees the true effect with a much smaller standard error (pooled \(n\) is 100 times larger) and can also detect publication bias by comparing the shape of published-vs-unpublished results. Single-study significance is fragile; meta-analysis is the right level of inference. (d) A power-\(0.5\) study is a coin flip for detection if the effect exists — and a headline conflates “detected” with “exists.” Given publication bias and the file-drawer problem, a single low-power “significant” result is weak evidence; wait for replication before believing the claim.

E.31 Chapter 31 — Sample Size Planning

Solutions to the odd-numbered exercises in Chapter 31 — Sample Size Planning.

StatLens Exercises

  1. (a) Slide \(n\) upward on the Power & Error Visualizer until the power readout crosses 0.80. This happens at about \(n \approx 40\) (the exact value depends on the tool’s tail conventions). (b) About \(n \approx 55\) achieves 90% power. Each additional percentage point of power costs more \(n\) than the last, because power approaches its ceiling of 1.0 asymptotically — the “last 10%” is always more expensive than any earlier 10%. (c) Much smaller — about \(n \approx 10\). Larger true effects are easier to detect, so a bigger \(\delta/\sigma\) dramatically reduces the \(n\) required for the same power. (d) The required \(n\) depends on the true effect size, which is unknown before the study. With no prior estimate, a researcher can: (i) use the smallest effect of practical or clinical interest, so the study is powered to detect anything worth caring about; (ii) run a pilot study first to estimate \(\sigma\) and a plausible \(\delta\); or (iii) fix \(n\) and instead report the minimum detectable effect at the planned \(n\).
  2. (a) \(n_{\text{per group}} = 2 \cdot (2.8 / 0.5)^2 = 2 \cdot (5.6)^2 = 62.7 \to 63\) per group. (b) The Power & Error Visualizer models a one-sample test, so its slider crosses 0.80 near \(n \approx 32\) — roughly half of 63. That is not a contradiction: a two-sample comparison needs about twice the one-sample \(n\) because the two-sample standard error carries a \(\sqrt{2}\) factor. The tool confirms the logic (power rises past 0.80), but read its \(n\) as the one-sample equivalent, not the per-group count for two samples. (c) Divide by \(1 - 0.20 = 0.80\): enroll \(63 / 0.80 \approx 79\) per group (round up to be safe). (d) At \(n = 30\), \(\delta/\sigma = 0.3\), \(\alpha = 0.05\) two-sided, power is only about \(0.21\) — a 21% chance of detecting a real small effect. Do not trust a non-result from this study; it was severely underpowered.