Math · AP Statistics ★★☆ Medium UNIT 3 OF 0

AP Statistics Unit 3: Collecting Data — Free Review Games.

This unit covers sampling methods, observational studies, experiments and bias — essential concepts for AP Statistics. Use our interactive study games to test your understanding, or review questions in traditional format below.

📋 200 questions ⏱ ~25 min 📊 12-15% of exam
Math Beast
Practice arena

Pick a mode. Play.

Answer questions as fast as you can. 2 minutes on the clock. Build streaks for bonus points!

Plain-text mode

Don't want to play?

All 200 questions below, each with the worked answer and a written explanation. Click any question to expand it.

Q1. A census collects data from:
A Every member of a population
B A sample
C Volunteers only
D Random groups

A census attempts to contact every individual in the population.

Q2. In a simple random sample (SRS), every:
A Group of size n has an equal chance of being selected
B Individual has the same opinion
C Variable is measured
D Sample is biased

In an SRS, every possible sample of size n is equally likely.

Q3. An observational study:
A Observes without imposing treatments
B Imposes treatments on subjects
C Always has a control group
D Uses double-blinding

Observational studies measure outcomes without intervening.

Q4. An experiment:
A Imposes treatments to observe effects
B Only observes existing conditions
C Cannot have a control group
D Uses convenience sampling

Experiments deliberately impose treatments to study cause and effect.

Q5. Voluntary response samples are often biased because:
A People with strong opinions are more likely to respond
B Everyone responds equally
C They are random
D They use stratification

Self-selected samples overrepresent people with strong feelings about the topic.

Q6. In stratified sampling, the population is divided into:
A Homogeneous groups (strata) then sampled from each
B Random clusters
C Pairs
D One large group

Stratification divides into similar groups and samples from every stratum.

Q7. Cluster sampling involves:
A Randomly selecting entire groups, then surveying all in chosen groups
B Selecting the best clusters
C Dividing by strata
D Using convenience

In cluster sampling, you randomly pick clusters and survey everyone in them.

Q8. A confounding variable:
A Is associated with both the explanatory and response variables
B Is the response variable
C Is always controlled
D Does not affect the study

Confounders are related to both variables, making it hard to determine causation.

Q9. The placebo effect occurs when:
A Subjects respond to a dummy treatment because they believe it is real
B The treatment actually works
C The control group gets medicine
D The experiment fails

The placebo effect is a psychological response to receiving a sham treatment.

Q10. Double-blind means:
A Neither subjects nor evaluators know who gets treatment
B Only subjects are blind
C Only researchers are blind
D Everyone knows the treatment

Double-blind prevents both subject and evaluator bias.

Q11. Random assignment in experiments ensures:
A Treatment groups are roughly equivalent
B Every person is in the study
C The sample is large
D Results are published

Random assignment balances known and unknown confounders across groups.

Q12. Undercoverage occurs when:
A Some groups in the population are left out of the sampling frame
B Too many people respond
C The sample is too large
D All groups are equally represented

Undercoverage means certain segments of the population have no chance of being sampled.

Q13. Which can establish causation?
A A well-designed randomized experiment
B An observational study
C A survey
D A census

Only randomized experiments with proper controls can establish cause and effect.

Q14. Response bias occurs when:
A Subjects give inaccurate answers due to question wording or social pressure
B Too few people respond
C The sample is too small
D Random assignment fails

Response bias results from misleading questions or subjects' desire to give socially acceptable answers.

Q15. The principle of replication in experiments means:
A Using enough subjects to reduce chance variation
B Repeating the entire experiment
C Having one subject per group
D Measuring twice

Replication uses many subjects so results are not due to chance alone.

Q16. In systematic random sampling, how are participants selected from the population?
A Every kth individual is selected after choosing a random starting point
B The population is divided into subgroups and individuals are randomly selected from each subgroup
C Entire groups are randomly chosen and all members of each selected group are included
D Individuals are selected based on their convenience and availability to the researcher

In systematic random sampling, a random starting point is chosen among the first k individuals, then every kth person thereafter is selected. This differs from stratified sampling (choice B), where individuals are drawn from within each subgroup, and cluster sampling (choice C), where entire groups are included wholesale. Convenience sampling (choice D) is not a probability method at all.

Q17. A convenience sample is one in which:
A Researchers select individuals who are easiest to reach or most readily available
B Participants are selected using a random number generator to ensure equal probability of selection
C The population is divided into groups of equal size before participants are randomly chosen
D Each member of the population is assigned a number and selected at equally spaced intervals

A convenience sample consists of individuals who are easy for the researcher to access, which often introduces bias because the readily available group may not represent the broader population. Choices B, C, and D describe probability-based sampling methods — simple random sampling, stratified sampling, and systematic sampling, respectively — all of which reduce bias through randomization.

Q18. In a designed experiment, 'experimental units' are best defined as:
A The individuals or objects to which treatments are randomly assigned and applied
B The outcome variables that the researcher measures at the end of the study
C The different conditions or procedures being compared across groups
D The researchers responsible for administering the treatments to participants

Experimental units are the entities that receive treatments — they may be people, animals, plants, or objects. The outcomes measured on these units are called response variables (choice B), the conditions being compared are the treatments (choice C), and the people running the study are investigators, not experimental units.

Q19. The primary purpose of including a control group in a designed experiment is to:
A Provide a baseline against which the treatment group's response can be compared
B Ensure that all participants are aware they are receiving the active treatment
C Eliminate every source of variability in the experimental results
D Increase the total number of participants and thereby reduce sampling error

A control group receives no treatment (or a placebo) and serves as a reference point. Without a control group, it is impossible to determine whether changes in the treatment group result from the treatment itself or from other factors such as the passage of time or the Hawthorne effect. Choice C is incorrect because no experimental design eliminates all variability — the goal is to reduce and account for it.

Q20. A clinical trial is described as single-blind. This means:
A Participants do not know which treatment they are receiving, but the researchers administering the study do
B Neither the participants nor the researchers know which treatment each participant is receiving
C Researchers do not know the hypothesis being tested during data collection
D Participants are prevented from communicating with one another about their treatment assignments

In a single-blind study, only the participants are kept unaware of their treatment assignment, while researchers retain that knowledge. This design reduces participant-side response bias but does not control for researcher bias (such as treating patients differently based on group assignment). When both participants and researchers are unaware, the study is double-blind (choice B), which provides stronger protection against bias.

Q21. A lurking variable is best described as:
A A variable associated with both the explanatory and response variables that is not included in the study
B The primary outcome variable whose change the researcher is trying to measure
C A variable directly manipulated by the researcher to observe its effect on the response
D Any variable measured during an experiment that turns out to be statistically insignificant

A lurking variable is not part of the analysis but can influence the relationship between the explanatory and response variables, potentially creating a spurious or misleading association. This is distinct from the response variable (choice B) and the explanatory or treatment variable (choice C). When a lurking variable is unaccounted for in a study, it becomes a confounding variable that can distort conclusions.

Q22. In a designed experiment, a 'treatment' refers to:
A A specific condition applied to experimental units, defined by one or more levels of the explanatory variable
B Any medical or therapeutic intervention, regardless of whether the study uses random assignment
C The measurement recorded on each experimental unit at the conclusion of the experiment
D The random procedure used to divide participants into groups before the study begins

A treatment is a specific experimental condition — for example, a particular drug dose, a teaching method, or a fertilizer type. The term applies broadly beyond medicine, making choice B too narrow. The measurement taken at the end is the response variable (choice C), and the process of dividing participants is random assignment (choice D), which is a method for allocating treatments, not the treatment itself.

Q23. A school administrator wants to survey student opinions about a new cafeteria menu. She randomly selects 6 homeroom classes and surveys every student in each of those classes. Which sampling method does this represent?
A Cluster sampling
B Stratified random sampling
C Systematic random sampling
D Simple random sampling

In cluster sampling, the population is divided into groups (clusters), some clusters are randomly selected, and every individual within those clusters is included. Here, homeroom classes serve as the clusters. In stratified random sampling (choice B), individuals would be randomly drawn from within each group — entire groups are never included wholesale. The key distinction is that cluster sampling selects groups, while stratified sampling samples from groups.

Q24. A magazine mails surveys to 2,000 subscribers asking about reading habits, but only 175 are returned and used to draw conclusions about all subscribers. The primary source of bias is:
A Non-response bias, because subscribers who return the survey may differ systematically from those who do not
B Undercoverage, because not all subscribers received a physical copy of the survey
C Voluntary response bias, because subscribers originally chose to purchase the magazine
D Response bias, because subscribers are likely to exaggerate their reading habits to please the publisher

Non-response bias occurs when those who respond differ meaningfully from those who do not. Highly engaged or strongly opinionated readers may be more likely to return the survey, making the 175 respondents unrepresentative of all 2,000 subscribers. This is distinct from voluntary response bias (choice C), which arises when individuals actively seek out a survey rather than being randomly selected first. Here, subscribers were selected first and then chose whether to respond.

Q25. A national researcher studying middle school students first randomly selects 15 school districts, then randomly selects 4 schools within each district, then randomly selects 30 students from each school. This procedure is best described as:
A Multistage sampling
B Stratified random sampling
C Cluster sampling
D Systematic random sampling

Multistage sampling applies random selection at multiple successive levels (districts, then schools, then students). This differs from simple cluster sampling (choice C), where all members of selected clusters are automatically included without further sampling within them. In stratified sampling (choice B), the population is divided into subgroups and individuals are randomly selected from each subgroup — the subgroups themselves are not sampled.

Q26. In a matched pairs experimental design, which of the following most accurately describes the procedure?
A Subjects are paired by shared characteristics, and one member of each pair is randomly assigned to each treatment
B All subjects receive both treatments in a fixed sequence based on their order of enrollment
C Subjects are randomly divided into two equal groups with no pairing or blocking applied
D Researchers match each subject with a statistically similar individual from a separate external population

In a matched pairs design, subjects with similar characteristics (such as age or baseline fitness) are paired, and then random assignment determines which member of each pair receives which treatment. This reduces variability attributable to the matching variable, improving sensitivity. Choice B lacks the essential element of random order assignment, which would introduce carryover or order effects. Choice C describes a completely randomized design with no pairing.

Q27. Which of the following statements about a completely randomized design is accurate?
A All experimental units are randomly assigned to treatment groups without any prior grouping or blocking
B Subjects are first grouped by a relevant characteristic before random assignment occurs within each group
C Every subject experiences all treatments, with the order of treatments randomly determined for each subject
D Treatments are assigned in a rotating sequence to guarantee exactly balanced group sizes in every replication

A completely randomized design uses random assignment as its only organizational structure — there is no blocking or pre-grouping. Choice B describes a randomized block design, which accounts for a known source of variability. Choice C describes a crossover or within-subjects design. A completely randomized design is simpler to implement but less efficient when a blocking variable is known to strongly influence the response.

Q28. A researcher studying whether a new fertilizer increases crop yield divides experimental plots by soil type and then randomly assigns plots within each soil type to either the fertilizer treatment or the control. The primary statistical advantage of blocking by soil type is:
A Reducing within-group variability by controlling for a known source of variation in crop yield
B Ensuring that all plots in the experiment have identical soil composition after blocking
C Preventing farmers from knowing which plots received the new fertilizer treatment
D Increasing the number of treatment groups to make the results more broadly generalizable

Blocking controls for a known source of variability so it does not obscure the treatment effect. By comparing fertilizer to no-fertilizer within plots of similar soil type, the researcher removes soil-type variability from the experimental error, making differences between treatments easier to detect. Blocking does not equalize soil composition (choice B) — it accounts for existing differences by grouping similar plots together before randomization.

Q29. A survey asks: 'Do you support the city council's reckless decision to slash funding for critical public safety services?' This question is most likely to produce biased responses due to:
A Leading question wording that steers respondents toward a particular answer
B Non-response bias from respondents unfamiliar with city council decisions
C Undercoverage of groups who do not follow local government news
D Voluntary response from only politically engaged citizens choosing to participate

The question uses emotionally loaded language ('reckless' and 'critical') that frames one position as the reasonable answer, pushing respondents toward a negative view of the council's decision. This is response bias caused by question wording. Non-response bias (choice B) and undercoverage (choice C) relate to who participates in the survey, not to how the question itself is phrased.

Q30. A quality control manager wants to assess the average weight of cereal boxes rolling off a production line. Why would a well-designed sample typically be preferred over a census of every box produced?
A A census is often impractical, time-consuming, or — in the case of destructive testing — physically impossible
B A sample always produces more accurate estimates of the population mean than measuring every unit
C Government regulations prohibit conducting a census for commercial manufacturing quality control
D Sampling completely eliminates measurement error, while a census introduces systematic errors

Practical constraints make a census impractical for large, continuous production runs. In destructive testing scenarios — such as testing battery life by running them until they fail — a census would destroy the entire inventory. Choice B is false: a census, if feasible, would have less sampling error than a sample. Choice D is false because sampling does not eliminate measurement error; it introduces sampling variability that a census avoids.

Q31. An experiment tests the effect of caffeine dose (0 mg, 100 mg, or 200 mg) and time of day (morning or afternoon) on reaction time. How many distinct treatments exist in this experiment?
A \(3 \times 2 = 6\) treatments
B \(3 + 2 = 5\) treatments
C 2 treatments, one for each time of day
D 3 treatments, one for each caffeine dose level

In a factorial experiment, the total number of treatments equals the product of the number of levels for each factor. With 3 caffeine levels and 2 time-of-day levels, there are \(3 \times 2 = 6\) distinct treatment combinations (e.g., 0 mg morning, 0 mg afternoon, 100 mg morning, 100 mg afternoon, 200 mg morning, 200 mg afternoon). Adding the levels (choice B) is a common error that fails to account for all possible combinations of the two factors.

Q32. A researcher surveys the first 60 customers entering a hardware store on a weekday morning to study home improvement spending habits among city residents. Which problem is MOST fundamental to this study's validity?
A Weekday morning hardware shoppers are not representative of city residents broadly, as they likely differ in employment status, income, and spending behavior
B A sample of 60 is too small to draw any valid conclusions about a city's population
C The researcher should have obtained written informed consent before beginning the survey
D Surveying customers in a single store creates response bias because they may overreport spending

The most fundamental flaw is that the convenience sample is drawn from a location and time that systematically excludes most city residents. Weekday morning hardware shoppers are likely to be retired individuals, contractors, or homeowners — not a cross-section of the city. Increasing the sample size (choice B) would not fix this bias; a larger convenience sample would remain equally unrepresentative.

Q33. An observational study finds that adults who drink coffee daily have lower rates of a certain liver condition compared to non-coffee drinkers. Which conclusion is most appropriate based on this finding?
A There is an association between daily coffee consumption and lower rates of the liver condition, but causation cannot be established from an observational study
B Daily coffee consumption causes a reduction in the risk of the liver condition based on the observed data
C Coffee drinkers are healthier overall, so the study should control for general health before drawing any conclusions
D The association is statistically meaningless because observational studies cannot detect real relationships between variables

Observational studies identify associations but cannot establish causation because subjects self-select into groups, making it impossible to rule out confounding variables. Coffee drinkers may differ from non-drinkers in diet, exercise habits, alcohol consumption, or other lifestyle factors that could explain the difference. Choice B incorrectly claims causation from observational data. Choice D is wrong — observational studies can and do detect genuine associations; the limitation is in causal inference, not detection.

Q34. A researcher blocks by income level when designing an experiment on tutoring program effectiveness. A colleague stratifies by income level when designing a survey of student performance. Which statement best distinguishes these two procedures?
A Blocking is used in experiments to reduce unexplained variability in the response; stratification is used in sampling to ensure proportional representation of subgroups
B Blocking and stratification are statistically identical procedures that differ only in the name applied depending on whether the context is experimental or observational
C Stratification applies only when subgroups have unequal sizes, while blocking is used only when subgroups are of equal size
D Blocking completely eliminates the effect of income on the response, whereas stratification only adjusts for income during analysis

Though conceptually analogous, blocking and stratification serve different purposes. In a designed experiment, blocking groups experimental units by a known variable before randomizing treatments within each block, thereby reducing variability that would otherwise obscure the treatment effect. In survey sampling, stratification divides the population into subgroups to ensure each is adequately represented in the sample. Choice B is incorrect — while related in spirit, the procedures operate in different contexts and with different statistical goals.

Q35. Data show that cities with more hospitals have higher overall death rates. A commentator concludes that hospitals are therefore dangerous and cause deaths. Which of the following best explains why this causal claim is not supported?
A Population illness burden is a confounding variable: sicker populations attract more hospital construction and also have higher mortality, creating a spurious association
B The relationship reflects response bias, because hospital administrators may overreport deaths to secure greater funding allocations
C Death rate should be treated as the explanatory variable rather than the response variable in this analysis
D A larger sample of cities would reverse the observed association, showing that hospitals actually reduce death rates

This is a classic confounding scenario. Population illness burden drives both the number of hospitals (more hospitals are built where illness is prevalent) and the death rate (sicker populations have higher mortality). Without accounting for this confounding variable, the association between hospitals and deaths is misleading and does not support a causal claim. Choice B misidentifies the problem as response bias, which involves misreporting of data rather than the influence of an unmeasured third variable.

Q36. A company recruits volunteers who specifically request a new energy drink and compares their monthly productivity scores to a group who did not request the drink. The most serious flaw preventing causal inference from this study is:
A Lack of random assignment: volunteers who seek out the energy drink likely differ from non-requesters in motivation, diet, or baseline productivity, confounding the comparison
B Lack of blinding: participants who know they are consuming the energy drink may experience a placebo effect that inflates productivity scores
C Lack of replication: the study needs to be repeated at multiple companies before any conclusion can be drawn
D Lack of a large enough sample: with more volunteers, random differences between groups would balance out

The most critical flaw is the absence of random assignment. People who actively seek out an energy drink likely differ in ways related to the outcome — they may be more motivated, health-conscious, or already accustomed to higher caffeine intake — creating systematic confounding. Without random assignment, observed productivity differences cannot be attributed to the drink itself. While blinding (choice B) is a genuine concern, it is secondary because the groups are not comparable to begin with.

Q37. A hospital compares success rates for two surgical procedures. Procedure A shows a 72% overall success rate versus 61% for Procedure B. However, when patients are separated into low-risk and high-risk groups, Procedure B has a higher success rate than Procedure A within both groups separately. Which of the following best explains this finding?
A Simpson's paradox: a confounding variable (patient risk level) reverses the direction of the association when the data are disaggregated into subgroups
B Non-response bias: high-risk patients who experienced failures were less likely to be included in the overall dataset
C Undercoverage: low-risk patients are underrepresented in the sample for Procedure B, distorting the aggregate comparison
D Replication failure: the results would align if the study were conducted at a different hospital with a different patient mix

Simpson's paradox occurs when a trend in aggregated data disappears or reverses when the data are divided into subgroups. Here, Procedure A was likely used more frequently on low-risk patients, who have higher baseline success rates regardless of procedure. This imbalance inflates Procedure A's aggregate rate. When controlling for risk level, Procedure B is superior in every subgroup. This illustrates why analyzing data only at the aggregate level can be deeply misleading when a confounding variable is unaccounted for.

Q38. Researchers randomly assign 120 adult volunteers with chronic back pain to receive either a new medication or a placebo for 12 weeks. The medication group shows statistically significantly greater pain reduction. Which of the following is the most accurate and complete conclusion?
A The experiment provides evidence that the medication caused pain reduction in this sample, but generalizability to all chronic back pain patients is limited because volunteers were used rather than a random sample
B Because participants were randomly assigned, results can be generalized to all adults with chronic back pain worldwide
C Because the study used volunteers, causation cannot be established regardless of the random assignment procedure
D The statistically significant result proves that every individual who takes this medication will experience pain reduction

Random assignment (not random sampling) is what justifies causal inference in an experiment — it ensures that treatment groups are comparable, so observed differences can be attributed to the treatment. However, scope of generalizability depends on how subjects were recruited. Since these are volunteers, not a random sample, results may not extend to all chronic back pain patients. Choice C is a key error: random assignment is sufficient for causation, even without random sampling. Choice B overclaims by extending results globally from a non-random sample.

Q39. A researcher wants to test whether a new reading intervention improves literacy scores among elementary students. She knows that parental education level is a strong predictor of literacy outcomes. Which experimental design would be most effective, and why?
A A blocked design grouping students by parental education level before randomly assigning within each block, reducing variability and improving the ability to detect the intervention's effect
B A completely randomized design, because random assignment alone guarantees perfect balance of parental education across treatment and control groups in every sample
C An observational study comparing students whose parents voluntarily choose the intervention to those who opt out, stratified by parental education during analysis
D A voluntary response design inviting families most interested in literacy improvement to self-select into the treatment group

When a variable is known to strongly predict the response, blocking on that variable before random assignment makes the design more sensitive by reducing within-block variability. Comparisons of treatment versus control are then made among students with similar parental education backgrounds, isolating the intervention's effect more cleanly. Choice B overstates the guarantee of random assignment — in small samples especially, random assignment may not perfectly balance all characteristics, and blocking is more reliable when a strong predictor is known. Choice C is an observational study and cannot establish causation.

Q40. A social media platform posts the following prompt: 'Should this platform ban all political advertising? Tap to vote!' Of 55,000 respondents, 79% say yes. Which combination of issues makes this result unreliable as an estimate of all platform users' opinions?
A Voluntary response bias and undercoverage: only users who feel strongly self-select into the poll, and social media users are not representative of all adults whose opinions might matter
B Non-response bias and response bias: low overall participation and social desirability pressure cause users to misrepresent their true opinions
C Undercoverage and question wording bias: the question is neutrally worded but excludes users who lack internet access
D Cluster sampling bias and non-response bias: platform users form natural demographic clusters that skew political preferences in one direction

This poll suffers from voluntary response bias — only users who feel strongly about political advertising bother to tap and vote, making the sample self-selected and unrepresentative of all users. Additionally, social media users as a population differ from the broader adult population in age, political engagement, and internet access, creating undercoverage if conclusions are extended beyond platform users. Choice B incorrectly frames the issue as low participation (55,000 responded, which is high) and conflates it with response bias, which involves misreporting rather than self-selection.

Q41. What is the primary purpose of randomly assigning subjects to treatment groups in an experiment?
A To ensure the sample is representative of the population
B To eliminate the placebo effect from the study
C To reduce the effect of confounding variables by distributing them evenly across treatment groups
D To guarantee that all individuals in the population have an equal chance of being selected for the study

Random assignment distributes potential confounding variables—both known and unknown—roughly equally across treatment groups, making the groups comparable before the treatment is applied. This is distinct from random selection, which addresses representativeness of the sample (choice A). Random assignment does not eliminate the placebo effect; blinding addresses that concern (choice B). Equal chance of selection describes random sampling, not random assignment (choice D).

Q42. A researcher divides a school's student population into freshmen, sophomores, juniors, and seniors, then randomly selects a fixed number of students from each group. Which sampling method is this?
A Cluster sampling
B Systematic random sampling
C Stratified random sampling
D Simple random sampling

Stratified random sampling involves dividing the population into subgroups (strata) based on a shared characteristic, then randomly sampling from each stratum. The strata here are the grade levels. Cluster sampling (choice A) involves dividing the population into groups and selecting entire groups, not random individuals from each group. Systematic sampling (choice B) involves selecting every \(k\)th individual from a list. Simple random sampling (choice D) gives every individual an equal chance of selection without dividing into subgroups first.

Q43. A telephone survey is conducted using only landline phone numbers to study the opinions of city residents. Which type of bias is most likely introduced?
A Nonresponse bias
B Response bias
C Undercoverage bias
D Voluntary response bias

Undercoverage bias occurs when some members of the population have little or no chance of being included in the sample. People without landlines—who may tend to be younger, lower-income, or more mobile—are systematically excluded from the sampling frame. Nonresponse bias (choice A) occurs when individuals chosen for the survey decline to participate. Response bias (choice B) occurs when participants give inaccurate answers. Voluntary response bias (choice D) occurs when individuals self-select to participate in a survey rather than being chosen by the researcher.

Q44. In a clinical trial testing a new pain medication, some participants receive a sugar pill that looks identical to the real medication. What is the primary statistical purpose of including this group?
A To give participants who might be harmed by the real medication a safer alternative
B To control for the psychological effect of believing one is receiving treatment
C To ensure the experiment has a sufficiently large sample size
D To allow researchers to compare different doses of the medication

A placebo controls for the placebo effect—the tendency for participants to report improvement simply because they believe they are receiving treatment. By comparing the active treatment group to a placebo group, researchers can isolate the actual pharmacological effect of the medication rather than the effect of expectation. The placebo group is not primarily a safety measure (choice A). Sample size is not the reason for including a placebo group (choice C). A dose comparison would require multiple active treatment groups, not a placebo (choice D).

Q45. Which of the following best distinguishes an experiment from an observational study?
A Experiments always use random sampling to select participants; observational studies do not
B In an experiment, the researcher actively imposes a treatment on subjects; in an observational study, the researcher does not intervene
C Observational studies are always conducted in natural settings; experiments occur only in laboratories
D Experiments measure quantitative variables; observational studies measure categorical variables

The defining feature of an experiment is that the researcher actively imposes a treatment or condition on subjects and observes the effect, rather than simply recording what occurs naturally. Choice A is incorrect because experiments do not require random sampling of participants—many experiments use volunteers. Choice C is incorrect because experiments can occur in natural field settings and observational studies can occur in controlled environments. Choice D is false; both study types can involve either type of variable.

Q46. A researcher mails surveys to 500 randomly selected households, but only 120 are returned. The researcher uses the 120 returned surveys to draw conclusions about the broader population. What is the primary statistical concern with this approach?
A The sample size of 120 is too small to produce any valid statistical conclusions
B The original mailing list was not a true random sample of all households
C Nonresponse bias may cause the 120 respondents to differ systematically from the 380 who did not respond
D Voluntary response bias invalidates any conclusion because only self-selected participants were mailed surveys

Nonresponse bias occurs when the individuals who respond differ systematically from those who do not. The 120 households that returned surveys may be more engaged, more opinionated, or have more time—making them unrepresentative of all 500 selected households and, in turn, the broader population. Sample size alone is not the fundamental problem (choice A); the quality and representativeness of responses matter more. Choice B is incorrect because the original selection was random. Voluntary response bias (choice D) refers to self-selected samples, not to nonresponse among a randomly chosen group.

Q47. In an experiment investigating whether listening to classical music improves focus, a group of participants completes the cognitive task in silence and receives no music treatment. What role does this group play in the experiment?
A Experimental group
B Placebo group
C Control group
D Stratified group

The control group receives no treatment (or the standard/baseline condition) and provides a comparison point for evaluating the effect of the treatment. Without a control group, researchers cannot determine whether any observed differences in focus are due to the music or to other factors. The experimental group (choice A) receives the treatment being tested—here, the music. A placebo group (choice B) receives an inactive substitute designed to mimic the treatment; silence is simply the absence of treatment. A stratified group (choice D) is not a standard experimental role.

Q48. A researcher wants to survey registered voters in a county and uses the county's official voter registration database to select participants. What does this database represent in the context of the study?
A The parameter of interest for the study
B The sampling frame from which the sample is drawn
C The sampling method used to select respondents
D The population to which the results will be generalized

A sampling frame is the list or set of individuals from which a sample is actually selected. Here, the voter registration database is the source from which the researcher draws the sample. If the database does not include all members of the target population (e.g., eligible voters who have not yet registered), undercoverage bias may result. The parameter of interest (choice A) is a numerical characteristic of the population, not a list. The sampling method (choice C) describes how individuals are chosen from the frame—for example, SRS or stratified sampling. The population (choice D) is the broader group the researcher wishes to study, which may be larger than the frame.

Q49. A quality control inspector wants to select 50 light bulbs from a production line that manufactures 5,000 bulbs per shift. She randomly selects a starting bulb between position 1 and 100, then inspects every 100th bulb after that. Which sampling method is she using, and what is a potential concern unique to this method?
A Stratified random sampling; the production batches may be unequal in size, causing imbalance
B Cluster sampling; entire production batches rather than individual bulbs are the sampling unit
C Simple random sampling; selecting every 100th bulb does not give all bulbs an equal probability of selection
D Systematic random sampling; a periodic pattern in the production process could align with the sampling interval and introduce bias

Selecting a random starting point and then every \(k\)th unit is systematic random sampling. Each unit does have an equal probability of selection (since the start is random), so choice C is incorrect. The critical concern unique to systematic sampling is periodicity: if the production process has a cycle whose period matches the sampling interval—for example, if the machine is recalibrated or a shift change occurs every 100 bulbs—then the sample could be biased toward particular production conditions. This does not apply to stratified or cluster sampling (choices A and B), which describe fundamentally different designs.

Q50. A study finds that neighborhoods with more fast food restaurants have higher rates of childhood obesity. A researcher concludes that the presence of fast food restaurants causes childhood obesity. Which of the following variables is most likely a confounding variable in this study?
A The total number of all restaurants (fast food and non-fast food combined) in the neighborhood
B The socioeconomic status of the neighborhood
C The number of children enrolled in schools within the neighborhood
D The geographic size of the neighborhood

A confounding variable is associated with both the explanatory variable and the response variable. Socioeconomic status (SES) is associated with fast food restaurant density (lower-income areas tend to have more fast food outlets and fewer grocery stores) and with childhood obesity rates (lower-income families may have reduced access to healthy food and safe recreational space). This means SES could explain the observed association between fast food density and obesity without any direct causal link. Total restaurant count (choice A) is less directly tied to obesity outcomes. School enrollment (choice C) and neighborhood size (choice D) are not strongly linked to both variables simultaneously.

Q51. A high school has 400 freshmen, 300 sophomores, 200 juniors, and 100 seniors. A researcher selects a proportionally stratified random sample of 100 students. How many juniors should be included in the sample?
A 10
B 20
C 25
D 30

In proportional stratified sampling, each stratum contributes to the sample in the same proportion as it appears in the population. The total population is \(400 + 300 + 200 + 100 = 1000\) students. Juniors make up \(\frac{200}{1000} = 0.20\), or 20%, of the population. The sample should therefore include \(100 \times 0.20 = 20\) juniors. This ensures that each grade level is represented proportionally, unlike simple random sampling, which could by chance over- or under-represent a grade level.

Q52. A researcher wants to study the daily internet usage habits of adults in a large city. She surveys people at a public library on weekday mornings. Which group is most likely to be underrepresented in her sample?
A Adults who are retired
B Adults who work full-time during weekday hours
C Adults who regularly visit public libraries
D Adults who live within walking distance of the library

Undercoverage occurs when certain groups in the population are systematically less likely to appear in the sample. Adults who work full-time are unlikely to be at a public library on weekday mornings, so they will be largely excluded from this convenience sample. This is a serious problem because full-time workers may have different internet usage patterns than those who are available during weekday mornings. Retired adults (choice A) actually have more weekday flexibility and may be overrepresented. Regular library visitors (choice C) and people living nearby (choice D) would be overrepresented, not underrepresented.

Q53. In a double-blind clinical trial testing a new antidepressant, neither participants nor the physicians evaluating them know which participants received the drug versus the placebo. Which of the following best explains why the physicians evaluating outcomes are also kept unaware of treatment assignments?
A To prevent physicians from accidentally prescribing additional medications to participants in the treatment group
B To ensure that all participants receive identical dosages of either the drug or the placebo
C To prevent physicians' expectations from unconsciously influencing how they assess and record participants' outcomes
D To protect participant confidentiality in compliance with medical privacy regulations

Blinding the evaluating physicians prevents experimenter bias—if a physician knows a patient received the active drug, they may (even unconsciously) interpret ambiguous symptoms as improvement or interact differently with that patient. Since depression assessment often involves subjective clinical judgment, this bias could distort the results in favor of the treatment. Choice A describes an administrative error-prevention measure, not a statistical rationale. Choice B addresses dosing consistency, which is handled by the trial protocol, not blinding. Choice D addresses ethics and privacy, not the statistical purpose of double-blinding.

Q54. A researcher testing whether a new instructional approach improves writing scores separates students into high-skill and low-skill groups based on their previous writing scores, then randomly assigns students within each group to receive either the new approach or the standard approach. Why does the researcher block by prior skill level?
A To ensure that all students in the study begin with the same writing ability
B To reduce variability within groups and increase the experiment's ability to detect a treatment effect
C To prevent the placebo effect from influencing how students respond to the instructional approach
D To guarantee that the student sample is representative of all students in the school

Blocking removes a known source of variability (prior skill level) from the unexplained error in the experiment, making it easier to detect the true effect of the new instructional approach. If students of widely varying skill levels were randomly mixed without blocking, the natural differences in scores could mask or inflate the treatment effect. Blocking does not equalize students' abilities (choice A)—it simply accounts for the differences that exist. Blocking is unrelated to the placebo effect (choice C), which is addressed through blinding. Representativeness (choice D) is a sampling concern, not a blocking rationale.

Q55. Researchers recruit 2,000 teenagers, assess their current dietary patterns and physical activity levels, and plan to follow them for 15 years, recording whether they develop cardiovascular disease. Which of the following best describes this study design?
A A retrospective observational study, because participants are recalling past behaviors
B A prospective observational study, because participants are followed forward in time from the present
C A randomized controlled experiment, because participants were recruited and assigned to study groups
D A matched pairs experiment, because each participant's baseline measurements are compared to their future outcomes

A prospective study recruits participants in the present, measures exposures or behaviors now, and then follows them forward in time to observe future outcomes. The researchers are not assigning any treatment—they are observing who naturally develops cardiovascular disease—making this an observational study, not an experiment (eliminating choices C and D). A retrospective study (choice A) looks backward in time, asking people who already have a disease to recall past behaviors. The key distinction here is the direction of time: forward (prospective) vs. backward (retrospective).

Q56. A county health department wants to survey households about access to primary care physicians. Rather than listing all households in the county individually, researchers randomly select 30 city blocks and then interview every household on each selected block. This is best described as which type of sampling?
A Stratified random sampling
B Systematic random sampling
C Cluster sampling
D Multistage sampling

In cluster sampling, the population is divided into groups (clusters)—here, city blocks—and entire clusters are randomly selected. All members of the chosen clusters are included in the sample. This is different from stratified sampling (choice A), where only a random subset of members is selected from each stratum, and every stratum is represented. Systematic sampling (choice B) involves selecting every \(k\)th individual from a list. Multistage sampling (choice D) would involve an additional stage—for example, randomly selecting blocks and then randomly selecting a subset of households within each selected block, rather than including all households.

Q57. A public health researcher is studying self-reported alcohol consumption. One group of participants completes an anonymous online questionnaire; another group is interviewed in person by a researcher in a clinical office setting. Compared to the anonymous questionnaire, which bias is the in-person interview most susceptible to?
A Undercoverage bias, because heavy drinkers are less likely to agree to in-person interviews
B Nonresponse bias, because scheduling in-person interviews leads to higher refusal rates
C Response bias due to social desirability, because participants may underreport alcohol intake to appear favorable to the interviewer
D Voluntary response bias, because in-person participants self-selected into the interview format

Response bias occurs when participants give inaccurate answers. In a face-to-face interview about a sensitive behavior like alcohol consumption, participants often underreport to avoid judgment—a social desirability effect. This effect is stronger when a human interviewer is present than in an anonymous questionnaire. Choice A describes undercoverage in who agrees to participate, not how they answer. Choice B describes nonresponse, which affects who participates, not the accuracy of responses given. Choice D describes voluntary response bias, which applies to self-selected samples, not to the format of a scheduled interview.

Q58. A university reports that its law school admits 42% of male applicants and 28% of female applicants overall. However, when admissions data are broken down by department, female applicants have a higher or equal acceptance rate in every individual department. Which statistical phenomenon best explains this apparent contradiction?
A Nonresponse bias caused by male applicants being more likely to complete the application process
B Response bias introduced by self-reported applicant qualifications on applications
C Simpson's paradox, where an association present in combined data reverses or disappears when the data are disaggregated
D Undercoverage bias because certain departments were not included in the department-level analysis

This is a classic example of Simpson's paradox: a trend that appears in aggregated data can reverse when the data are broken into subgroups. The explanation is that female applicants disproportionately apply to more competitive departments (with low acceptance rates for everyone). Even though women perform as well as or better than men within each department, their overall rate is lower because they are concentrated in harder-to-enter programs. The paradox arises from the confounding effect of department selectivity on the aggregate comparison. This is not a data collection or measurement problem (choices A, B, D)—it is a structural property of combining groups with different sizes and base rates.

Q59. A pharmaceutical company randomly assigns 200 patients to receive either a new antidepressant or a placebo. The physicians treating the patients know which treatment each patient is receiving. Patients do not know. Physicians assess improvement using subjective clinical interviews. Which flaw most seriously threatens the internal validity of this study?
A The sample of 200 patients is insufficient to detect a meaningful effect of an antidepressant
B Physicians who know the treatment assignment may unconsciously rate improvement differently for treated versus placebo patients, introducing experimenter bias
C Patients who receive the placebo may be harmed by not receiving active treatment, which could distort the comparison
D Random assignment cannot guarantee that the treatment and control groups are perfectly balanced on all baseline characteristics

Because physicians know who received the active drug and are conducting subjective clinical interviews, their expectations may unconsciously influence how they probe for symptoms, interpret ambiguous responses, or rate improvement. This experimenter bias directly threatens internal validity—differences between groups may reflect the physician's beliefs rather than the drug's effect. This is why double-blinding is the gold standard. Choice A is incorrect; 200 patients can be adequate depending on the expected effect size. Choice D is technically true but random assignment with 200 patients generally produces comparable groups; this is a minor concern relative to the active experimenter bias introduced by single-blinding. Choice C is an ethical concern, not a threat to internal validity.

Q60. A study of elementary school children finds a strong positive correlation between shoe size and reading ability. A principal concludes that children with larger feet have a natural advantage in reading. Which of the following is the most accurate critique of this conclusion?
A The correlation coefficient is not an appropriate measure for data collected from children
B Age is a lurking variable: older children tend to have both larger feet and stronger reading skills, creating a spurious correlation between shoe size and reading ability
C The study should have used an experiment rather than an observational study, because only experiments can reveal correlations
D The sample of elementary school children is not representative of all children, so the correlation cannot be trusted

A lurking variable is an unmeasured variable that drives a spurious association between two other variables. Age causes both foot growth and reading development. As children get older, both shoe size and reading ability increase—the two variables are correlated because they share a common cause (age), not because one causes the other. This is a classic example of a confounding lurking variable. Choice A is incorrect; correlation is statistically valid here, but the causal interpretation is wrong. Choice C is incorrect; observational studies routinely reveal correlations—experiments test causal hypotheses. Choice D addresses external validity (generalizability), not the flawed causal inference from foot size to reading ability.

Q61. A researcher designs a \(2 \times 2\) factorial experiment testing the effects of study time (1 hour vs. 3 hours) and study method (flashcards vs. practice problems) on exam scores. After analysis, the researcher finds that practice problems outperform flashcards only when combined with 3 hours of study, but the two methods produce similar results at 1 hour of study. Which of the following best describes this finding?
A A confounding effect of study time on the comparison of study methods
B An interaction effect between study time and study method
C A blocking effect that reduced variability in exam scores across groups
D A lurking variable effect introduced by students self-selecting their preferred method

An interaction effect occurs when the effect of one factor on the response variable depends on the level of a second factor. Here, the benefit of practice problems over flashcards is present at 3 hours but not at 1 hour—meaning the effect of study method depends on study time. This is the definition of an interaction. A confounding effect (choice A) would mean study time is associated with study method in a way that distorts the comparison, not that the effect of method varies across time levels. Blocking (choice C) is a design technique for controlling variation, not a description of a pattern in results. A lurking variable (choice D) was not introduced here because the study used random assignment to all four treatment combinations.

Q62. A local newspaper posts an online poll asking readers: 'Should the city raise property taxes to fund new public schools?' The poll receives 4,500 responses. A city council member cites the poll results as evidence of residents' true preferences. Which of the following most accurately describes the primary limitation of this poll?
A The sample size of 4,500 is insufficient to estimate preferences across the entire city
B Voluntary response bias causes the sample to over-represent residents with strong opinions, likely making results unrepresentative of the broader population
C The question wording is neutral enough that response bias is not a concern in this poll
D The poll should have stratified respondents by homeownership status to ensure valid conclusions

The fundamental flaw is voluntary response bias. People who choose to click on a poll and submit a response are self-selected—they tend to have stronger opinions (positive or negative) than the general public. Residents who are indifferent or only mildly opinionated are unlikely to participate. This makes the results systematically unrepresentative regardless of the large sample size, so choice A is incorrect. Choice C is plausible as written but sidesteps the main problem. Choice D suggests a design improvement that would not address voluntary response bias, since the self-selection problem persists even if respondents were stratified after the fact.

Q63. A researcher has 80 volunteers with insomnia and wants to test whether cognitive behavioral therapy (CBT) reduces sleep onset time compared to a sleep hygiene pamphlet. She suspects that baseline insomnia severity may affect how well each treatment works. Which experimental design best addresses this concern?
A Randomly assign all 80 participants to CBT or the pamphlet without accounting for severity, since randomization will tend to balance this variable
B Use only participants with the same baseline severity to eliminate this variable from the study
C Block by baseline insomnia severity, then randomly assign participants within each severity block to CBT or the pamphlet
D Conduct a matched pairs design by pairing each CBT participant with a pamphlet participant who has the same severity, selected after the study is complete

Blocking by baseline insomnia severity is the most powerful and appropriate design when a variable is known to affect the response. By randomly assigning within each severity block, the researcher ensures both treatment groups contain comparable proportions of mild, moderate, and severe cases, while also being able to estimate and remove the variability due to severity from the error term. This increases the power to detect a treatment effect. Choice A is statistically valid but less efficient with only 80 participants, where random chance imbalance is more likely. Choice B restricts generalizability and wastes data. Choice D describes retrospective matching after data collection, which does not achieve the same control as prospective blocking and can introduce selection bias.

Q64. To estimate average reading scores of fourth graders nationwide, researchers randomly select 80 school districts from a national list, then randomly select 4 schools within each chosen district, and finally randomly select 25 students from each chosen school. This is an example of which design, and what is its primary statistical disadvantage compared to a simple random sample of the same total size?
A Cluster sampling; entire districts rather than individual students are the ultimate sampling unit, so students outside selected districts have no chance of inclusion
B Stratified random sampling; districts may differ in size, causing students from smaller districts to be over-represented
C Multistage sampling; students within the same school tend to be more similar to each other than to students nationwide, reducing estimation precision relative to an SRS of the same size
D Systematic random sampling; selecting every \(k\)th district from a ranked list introduces periodic bias

This is multistage sampling: random selection occurs at three successive levels (district, school, student). The key statistical disadvantage is clustering: students within the same school share teachers, curricula, and community resources, making their scores more homogeneous than scores from students randomly selected across the entire country. This within-cluster similarity reduces the effective sample size and yields less precise estimates than a true SRS of \(80 \times 4 \times 25 = 8000\) students. However, multistage sampling is vastly more logistically feasible. Choice A describes single-stage cluster sampling (selecting whole districts with no further subsampling). Choice B is incorrect; stratification is not occurring here. Choice D describes systematic sampling, which is not the design used.

Q65. A large longitudinal observational study of 60,000 adults finds that regular exercisers have significantly lower rates of depression than sedentary adults, even after statistically adjusting for age, income, and social support. A health journalist writes: 'New study proves exercise cures depression.' Which of the following is the most precise and complete critique of this headline?
A The sample size of 60,000 is so large that even trivial effects become statistically significant, so the result has no practical meaning
B Observational studies cannot establish causation even with statistical controls, because unmeasured confounders may explain the association and reverse causation is possible—depression itself may reduce the likelihood of exercising
C The study should have used a matched pairs design to properly compare exercisers and sedentary individuals before concluding any relationship exists
D Statistical significance only addresses sampling variability and does not indicate whether the reduction in depression rates is large enough to be clinically important

Even a large, well-controlled observational study cannot prove causation. First, unmeasured confounders—variables not recorded in the study, such as personality traits, genetic predispositions, or chronic illness—may explain the association. Second, reverse causation is a serious concern: people experiencing depression may exercise less because of their condition, meaning the causal arrow may point from depression to inactivity rather than the reverse. Only a randomized experiment, where participants are assigned to exercise or not, can provide strong evidence for causation. Choice A is incorrect; large samples improve reliability, not distort significance. Choice C suggests a design improvement but does not identify why the causal claim is invalid from an observational study. Choice D raises a valid distinction between statistical and clinical significance, but does not address the word 'proves' or the causal language in the headline as directly as choice B.

Q66. In a simple random sample (SRS) of size \(n\) from a population, which property must hold?
A Every individual has an equal probability of being selected
B The population is divided into strata before sampling begins
C Individuals are selected at regular intervals from an ordered list
D Every possible sample of size \(n\) has an equal probability of being selected

An SRS requires that every possible sample of size \(n\) has an equal chance of being selected — a stronger condition than simply giving each individual an equal probability of inclusion. For example, in systematic sampling each individual may have equal probability of being chosen, but not every sample of size \(n\) is equally likely (since only samples consistent with the fixed interval can occur). Choice A describes a weaker requirement satisfied by several non-SRS methods. Choice B describes stratified sampling. Choice C describes systematic sampling.

Q67. A radio station asks listeners to call in and vote on whether a proposed city curfew for teenagers is a good idea. What type of sampling method does this represent?
A Systematic random sample
B Stratified random sample
C Voluntary response sample
D Cluster sample

A voluntary response sample consists of individuals who choose to participate on their own, typically those who feel most strongly about the issue. This produces biased results because respondents are self-selected — listeners who are especially opposed to or enthusiastic about the curfew are far more likely to call in. Choice A involves selecting every \(k\)-th individual from a list. Choice B divides the population into subgroups and samples from each. Choice D involves randomly selecting entire groups and surveying all their members.

Q68. The fundamental difference between an observational study and a controlled experiment is that in an experiment, the researcher:
A Always uses a larger sample size than in observational studies
B Collects data over a longer period of time
C Uses a randomly selected sample from the target population
D Deliberately assigns individuals to treatment groups

In an experiment, the researcher actively imposes treatments on subjects by assigning them to groups, which is what permits causal conclusions to be drawn. In an observational study, the researcher merely observes without controlling assignments. Choice A is false — observational studies can be larger than experiments. Choice B is false — time frame varies independently of study type. Choice C describes how a sample is obtained, which is desirable in both study types and does not distinguish an experiment from an observational study.

Q69. In a randomized clinical trial, a placebo is given to the control group primarily to:
A Control for the psychological effect of believing one is receiving a treatment
B Reduce the overall cost of running the experiment
C Ensure the experimental group receives a higher dose of the medication
D Eliminate the need for randomization between groups

The placebo controls for the placebo effect — the tendency for people to feel or report improvement simply because they believe they are receiving treatment. Without a placebo, any improvement in the treatment group could be partly due to this psychological effect rather than the medication itself. A placebo does not reduce costs (Choice B), affect the experimental group's dosage (Choice C), or substitute for randomization (Choice D), which serves the separate purpose of balancing known and unknown confounding variables across groups.

Q70. A biologist tests whether a new fertilizer increases plant growth. Some plants receive the fertilizer; others are grown under identical conditions but receive no fertilizer. The plants that receive no fertilizer are called:
A The control group
B The placebo group
C The experimental units
D The blocking group

The control group receives no treatment (or a standard treatment) and provides a baseline against which the fertilizer's effect can be measured. Choice B (placebo group) typically refers to a group receiving an inert treatment designed to mimic the real one — most relevant in human studies where psychological expectation matters. Choice C (experimental units) refers to all individual subjects under study, not just one group. Choice D (blocking group) is not a standard term for this role in experimental design.

Q71. A quality control inspector at a bottling factory examines every \(15\)th bottle coming off the production line. This is an example of:
A Systematic random sampling
B Simple random sampling
C Stratified random sampling
D Cluster sampling

Systematic random sampling involves selecting every \(k\)-th unit from a list or sequence, often after a randomly chosen starting point. Here, every \(15\)th bottle is inspected. Choice B (SRS) would require randomly selecting individual bottles from all bottles produced. Choice C (stratified) would divide bottles into distinct groups and sample from each group separately. Choice D (cluster) would involve randomly selecting entire production runs or batches and inspecting all bottles within the selected batches.

Q72. In a well-designed experiment, the principle of replication means that:
A Each treatment is applied to multiple experimental units
B The experiment is reproduced in multiple laboratories
C The same subject receives each treatment more than once
D The experiment includes both a treatment group and a control group

Replication means applying each treatment to multiple experimental units so that natural variability among units can be estimated and random chance is less likely to account for any observed differences. Without replication — if only one unit receives each treatment — it is impossible to distinguish a true treatment effect from random variability. Choice B describes scientific reproducibility across labs, a separate concept. Choice C describes the matched pairs design, not replication. Choice D describes the use of a control group, which is a related but distinct design principle.

Q73. A city public health office wants to estimate the proportion of residents who exercise regularly. They divide the city into 12 neighborhoods, randomly select 3 of those neighborhoods, and then survey every resident in each of the 3 selected neighborhoods. This is an example of:
A Multi-stage random sampling
B Stratified random sampling
C Cluster sampling
D Systematic random sampling

In cluster sampling, the population is divided into groups (clusters), clusters are randomly selected, and every member of each selected cluster is surveyed. Here, neighborhoods are the clusters, 3 are randomly chosen, and all residents in those neighborhoods participate. Choice A (multi-stage) would involve randomly selecting a subset of residents within each chosen neighborhood rather than surveying everyone. Choice B (stratified) would require randomly selecting individuals from every neighborhood. Choice D uses a fixed selection interval from an ordered list.

Q74. A study finds that children who own more books score higher on standardized reading tests. A researcher concludes that owning more books causes better reading ability. Which of the following is the most plausible confounding variable that could explain this association?
A The number of hours children spend watching television each week
B The age at which children first learn to read
C The geographic region in which children live
D Household income, which is associated with both greater book ownership and access to broader educational resources

A confounding variable is associated with both the explanatory variable (book ownership) and the response variable (reading scores), making it difficult to isolate book ownership as the cause. Household income affects both: wealthier families can afford more books and also provide tutoring, better schools, and enrichment activities that boost reading scores. Choice A (TV watching) may correlate with reading time but is not as directly tied to book ownership itself. Choices B and C are not strongly or systematically linked to both variables simultaneously.

Q75. A researcher stands outside a fitness center and surveys the first 60 people who exit, asking how many hours per week they exercise, in order to estimate the average exercise time for all adults in the city. The estimate from this sample is most likely to:
A Be accurate because the researcher had no control over who exited first
B Overestimate average exercise time because fitness center patrons exercise more than the general population
C Underestimate average exercise time because frequent exercisers are too tired to stop and answer
D Be unbiased because 60 is a sufficiently large sample

This is a convenience sample — whoever was easiest to reach was selected. People who go to a fitness center exercise far more than the typical adult, so the sample differs systematically from the general population, producing an overestimate. Choice A is incorrect — random order of exit from the fitness center does not correct for the fact that all respondents are fitness center patrons. Choice C is implausible and directionally wrong. Choice D is wrong because sample size alone cannot eliminate systematic bias; a biased sample of 60 is still biased.

Q76. Researchers are testing three study strategies — flashcards, practice tests, and re-reading — on student retention. They suspect that prior GPA may affect both study habits and retention outcomes. They divide students into high-GPA and low-GPA groups, then randomly assign students within each group to one of the three strategies. This design is a:
A Completely randomized design
B Matched pairs design
C Multi-stage sampling design
D Randomized block design

A randomized block design groups experimental units into blocks based on a variable expected to affect the response (here, GPA level), then randomly assigns treatments within each block. This controls for the effect of GPA and reduces within-group variability, making it easier to detect differences among study strategies. Choice A (completely randomized) would assign all students to strategies without accounting for GPA differences. Choice B (matched pairs) pairs individual subjects rather than grouping them into blocks. Choice C is a sampling method, not an experimental design.

Q77. A survey question reads: 'Do you support wasteful government spending on foreign aid that benefits other countries at the expense of American taxpayers?' Compared to a neutrally worded question on the same topic, responses to this question will most likely:
A Show a higher proportion opposing foreign aid due to the loaded and leading language
B Show a higher proportion supporting foreign aid due to social desirability effects
C Be unaffected because respondents are free to choose any answer option
D Show greater variability in responses compared to a neutral question

The question uses loaded, negatively framed language ('wasteful,' 'at the expense of American taxpayers') that steers respondents toward opposition — a form of response bias called a leading question or wording effect. This produces a systematic overestimate of opposition to foreign aid. Choice B (social desirability) describes people giving the answer they think is socially acceptable, which would likely favor supporting aid rather than opposing it. Choice C is incorrect — question wording substantially influences responses. Choice D is incorrect — the effect shifts the distribution in one direction rather than increasing spread.

Q78. A city council mails a survey to a random sample of 1,000 residents asking whether they support a new recycling ordinance. Of the 1,000 surveys mailed, 210 are returned. The council reports that 82% of respondents support the ordinance. Which of the following most seriously threatens the validity of the council's conclusion?
A A sample of 1,000 is too small to represent a large city
B Mail surveys always produce more biased results than phone surveys
C 82% approval should be converted to a decimal before reporting any statistics
D Residents who chose to respond may differ systematically from non-respondents, introducing nonresponse bias

With only a 21% response rate, nonresponse bias is the central concern. Residents who feel strongly enough to return the survey — perhaps those passionate about environmental issues or those strongly opposed — may not represent all residents. The 82% approval figure applies only to those who responded, and generalizing it to all 1,000 sampled residents (or to the entire city) is unjustified. Choice A is incorrect — 1,000 is a reasonable sample size. Choice B is a false generalization. Choice C is a trivial formatting issue with no bearing on statistical validity.

Q79. To study whether caffeine improves reaction time, researchers recruit 30 volunteers. Each volunteer completes a standardized reaction-time test twice — once after consuming a caffeinated drink and once after consuming a decaffeinated drink — with the order randomly assigned to each participant. This study uses:
A A completely randomized design
B A randomized block design with caffeine as the blocking variable
C A convenience sample without experimental controls
D A matched pairs design

A matched pairs design is a special form of blocking in which each subject serves as their own control by receiving both treatments. Because each person's baseline reaction time is controlled for, individual differences are removed as a source of variability, increasing the power of the study. The random ordering of conditions prevents order effects (e.g., fatigue or practice). Choice A would randomly assign different volunteers to the caffeinated or decaffeinated condition, with each person receiving only one treatment. Choice B uses blocks of multiple subjects, not individual subjects as their own controls. Choice C mischaracterizes the study.

Q80. Data collected across many cities shows a strong positive correlation (\(r = 0.91\)) between the number of hospitals per capita and the city's death rate per capita. A journalist concludes that hospitals must cause deaths and proposes that cities close hospitals to save lives. Which of the following best identifies the flaw in this reasoning?
A The study should have used a controlled experiment rather than computing a correlation coefficient
B Population health status is a lurking variable — cities with sicker and older populations need more hospitals and also have higher death rates
C A correlation of \(r = 0.91\) is not strong enough to draw meaningful conclusions
D The researcher should have reported \(r^2\) rather than \(r\) to measure the strength of the association

A lurking variable is one not included in the analysis that explains an observed association. Cities with older or sicker populations naturally require more hospitals and also experience higher death rates — the lurking variable (population health and age structure) drives both. This is a classic example of confusing correlation with causation. Choice A correctly notes that experiments support causal claims, but it does not identify the specific flaw in the journalist's argument. Choice C is wrong — \(r = 0.91\) is a very strong correlation. Choice D is a separate statistical consideration and not the core error.

Q81. A school district mails surveys to a random sample of 800 families asking about satisfaction with the district's technology programs. Of the 800 surveys mailed, 175 are returned. The district reports that 88% of families are satisfied with the programs. What is the primary threat to the validity of this conclusion?
A A sample of 800 families is too small to draw valid conclusions for a school district
B Families who chose to respond may be more engaged with school programs than non-respondents, creating nonresponse bias
C Satisfaction surveys should always be administered in person to prevent bias
D The district should have conducted a census rather than a sample to get accurate results

With only a 22% response rate, nonresponse bias is the central concern. Families who returned the survey likely differ from those who did not — they may be more engaged with the school, have stronger opinions about technology programs, or have had notably positive or negative experiences. Reporting 88% satisfaction as representative of all sampled families (or the entire district) is unjustified. Choice A is incorrect — 800 is a reasonable sample size. Choice C is a false generalization. Choice D may not be feasible and would not necessarily prevent nonresponse bias even in a census if some families decline to participate.

Q82. To study the dietary habits of high school students nationwide, researchers first randomly select 40 high schools from across the country, then within each selected school randomly select 25 students to participate in the study. This is an example of:
A Cluster sampling
B Stratified random sampling
C Multi-stage random sampling
D Systematic random sampling

Multi-stage sampling involves multiple rounds of random selection. Here, schools are randomly selected first (stage 1), then students within each school are randomly selected (stage 2). This differs from cluster sampling (Choice A), in which every member of each selected cluster is surveyed rather than a random subset. Choice B (stratified) would divide students into meaningful groups (strata) and sample from all of them. Choice D uses a fixed interval from an ordered list rather than separate random selections at each stage.

Q83. A hospital researcher wants to compare patient satisfaction across three departments: cardiology, oncology, and orthopedics. She randomly selects 25 patients from each department and surveys them. Which of the following best justifies using stratified random sampling rather than a simple random sample for this study?
A Stratified sampling is computationally simpler and faster than simple random sampling
B Stratified sampling guarantees representation from each department and can reduce overall sampling variability
C Stratified sampling ensures that each department's sample size reflects its proportion of total hospital patients
D Stratified sampling allows the researcher to survey all patients in any one department

Stratified random sampling divides the population into subgroups (strata) and samples from each, ensuring that each department is represented even if one department is much smaller than others. It also tends to reduce overall sampling variability by creating more homogeneous groups within each stratum. A simple random sample might by chance severely underrepresent a small department. Choice A is false — stratification adds logistical complexity compared to SRS. Choice C describes proportional stratified sampling, a specific type not implied by this design. Choice D is incorrect — stratified sampling takes a random sample from each stratum, not a census.

Q84. A researcher investigates the effect of sleep duration on reaction time and suspects that age may influence both sleep patterns and reaction time. She recruits volunteers from two age groups (20–30 years and 50–65 years), then randomly assigns participants within each age group to one of three sleep conditions: 5, 7, or 9 hours per night. What is the primary statistical advantage of blocking on age group in this design?
A Blocking on age removes the effect of age on reaction time entirely from the study
B Blocking on age allows the researcher to draw causal conclusions about the effect of age on reaction time
C Blocking on age eliminates the need for random assignment within each age group
D Blocking on age reduces variability within treatment groups, increasing the ability to detect a true effect of sleep duration

Blocking on a known source of variability keeps that source constant within each block, reducing within-group noise and improving the experiment's power to detect differences due to sleep duration. If age groups were mixed without blocking, age-related variability in reaction time would inflate the error term, making it harder to detect sleep effects. Choice A is incorrect — blocking accounts for age's influence but does not remove it from reality. Choice B is wrong — participants were not randomly assigned to age groups, so no causal claim about age is warranted. Choice C is wrong — random assignment within blocks is essential to the design and cannot be eliminated.

Q85. A researcher conducts a randomized experiment using 120 volunteers from a single university campus to test whether a mobile app reduces anxiety. She randomly assigns 60 volunteers to use the app for six weeks and 60 to use a control app. She finds a statistically significant reduction in anxiety scores for the app group (\(p = 0.003\)). Which of the following conclusions is most justified?
A The app causes anxiety reduction in all adults because the experiment used random assignment
B There is convincing evidence that the app caused anxiety reduction in this study's participants, though results may not generalize to all adults
C Because the sample was drawn randomly, results can be generalized to all college students
D A result of \(p = 0.003\) proves the app is clinically effective for everyone who uses it

Random assignment justifies causal conclusions — the app (not some other variable) caused the observed reduction in anxiety among participants. However, because the sample consists of volunteers from a single university campus, it likely does not represent all adults or even all college students, which limits external validity. Choice A overgeneralizes to all adults without justification. Choice C incorrectly conflates random assignment (which supports causation) with random sampling (which supports generalizability) — these are separate design features. Choice D confuses statistical significance with clinical efficacy or universal applicability.

Q86. To assess whether a new tutoring method improves student performance, a researcher randomly assigns 60 students to either the new tutoring method or a traditional method. All sessions are held in a standardized laboratory setting with identical materials and strict time controls. A colleague argues the study has poor external validity. Which of the following best supports this claim?
A The students were not randomly assigned to conditions, so causation cannot be established
B The study lacks a no-treatment control group, making comparisons invalid
C The artificial laboratory conditions may not reflect real classroom environments, limiting how well results generalize to typical educational settings
D A sample of 60 students is insufficient to establish statistical significance in an educational study

External validity refers to how well study results generalize to other settings, populations, and conditions. A tightly controlled laboratory environment may produce results that do not replicate in actual schools where distractions, varying materials, instructor differences, and student motivation differ considerably. The colleague is correctly questioning whether lab-based findings will hold in the real world. Choice A is incorrect — the question states random assignment was used. Choice B is incorrect — comparing two active instructional methods without a no-treatment group is a valid design for this research question. Choice D confuses statistical power with validity.

Q87. In a \(2 \times 2\) factorial experiment, researchers test the effects of exercise intensity (low vs. high) and diet type (standard vs. low-carbohydrate) on weight loss over 12 weeks. They find that the low-carbohydrate diet produces substantially greater weight loss than the standard diet among participants in the high-intensity exercise condition, but produces nearly identical weight loss to the standard diet among participants in the low-intensity exercise condition. This pattern of results most clearly indicates:
A A main effect of exercise intensity but no main effect of diet type
B A main effect of diet type but no main effect of exercise intensity
C An interaction between exercise intensity and diet type
D That neither factor has a meaningful independent effect on weight loss

An interaction effect occurs when the effect of one factor on the response variable depends on the level of another factor. Here, the benefit of the low-carbohydrate diet appears only when exercise intensity is high — the diet's effect changes depending on exercise level. This is the definition of an interaction. Choice A (main effect of exercise only) would imply exercise affects weight loss consistently across diet types without diet mattering differently at each exercise level. Choice B (main effect of diet only) would imply diet affects weight loss consistently across exercise levels. Choice D is contradicted by the described pattern, which shows a clear conditional relationship.

Q88. A researcher tests whether a new anti-inflammatory drug reduces knee pain more effectively than the current standard medication. Patients are randomly assigned to receive either the new drug or the standard drug, and neither the patients nor the physicians evaluating pain scores know which drug each patient received. Which of the following best explains why blinding the physicians — not just the patients — is important in this study?
A Blinding physicians ensures that the two treatment groups contain equal numbers of patients
B Blinding physicians prevents the placebo effect from influencing patient self-reports
C Blinding physicians prevents their outcome assessments from being influenced by expectations about which treatment is superior
D Blinding physicians eliminates the need for random assignment to balance confounding variables

Blinding physicians prevents observer bias (also called experimenter bias): a physician who knows a patient is receiving the new drug and expects it to be superior may unconsciously rate that patient's pain as lower, artificially inflating the apparent treatment benefit. This is why outcome assessors must also be blinded, not just patients. Choice A is incorrect — group size equality comes from randomization, not blinding. Choice B is incorrect — the placebo effect refers to changes in patients' own responses due to belief; blinding physicians addresses how evaluators measure outcomes. Choice D is incorrect — blinding and randomization serve distinct, complementary purposes and neither replaces the other.

Q89. A school district introduces a new reading curriculum in September and measures students' reading scores before the school year (mean = \(72\)) and after (mean = \(81\)), concluding that the curriculum caused the \(9\)-point improvement. A statistician argues this conclusion is flawed due to a threat to internal validity. Which of the following most likely explains what the statistician means?
A Students naturally improve in reading over the course of a school year regardless of curriculum, making it impossible to attribute the gain to the new curriculum without a control group
B A \(9\)-point improvement is too small to be considered statistically meaningful in educational research
C Reading scores should be measured at multiple points during the year rather than only before and after
D Because all students received the same curriculum, the study violates the principle of replication

Without a control group receiving the existing curriculum, the researcher cannot determine whether the \(9\)-point improvement resulted from the new curriculum or from natural student growth and maturation over the school year — a threat to internal validity called the history or maturation effect. Students typically improve in reading each year regardless of instructional approach, so the gain may have occurred without any curriculum change. A concurrent control group would allow the researcher to separate curriculum effects from natural growth. Choice B is an empirical claim requiring context — it is not the identified flaw. Choice C describes a measurement preference, not a validity threat. Choice D misapplies the replication principle.

Q90. A \(2 \times 2\) factorial experiment tests Factor A (levels \(A_1\) and \(A_2\)) and Factor B (levels \(B_1\) and \(B_2\)) on a quantitative response. The four cell means are: \((A_1, B_1) = 20\), \((A_1, B_2) = 40\), \((A_2, B_1) = 40\), \((A_2, B_2) = 20\). The marginal mean for \(A_1\) equals the marginal mean for \(A_2\) (both \(= 30\)), and the marginal mean for \(B_1\) equals the marginal mean for \(B_2\) (both \(= 30\)). Based solely on these means, which conclusion is most appropriate?
A There are main effects of both Factor A and Factor B, but no interaction between them
B There is a main effect of Factor A only, with no main effect of Factor B
C There are no main effects of either factor, but there is an interaction between Factor A and Factor B
D There is a main effect of Factor B only, with no main effect of Factor A

Main effects are assessed by comparing marginal means. Since \(A_1\) and \(A_2\) both have marginal means of \(30\), there is no main effect of Factor A. Since \(B_1\) and \(B_2\) both have marginal means of \(30\), there is no main effect of Factor B. However, the cell means reveal a striking interaction: when Factor A is at \(A_1\), Factor B at \(B_2\) yields a higher response (\(40\) vs. \(20\)); when Factor A is at \(A_2\), Factor B at \(B_1\) yields a higher response (\(40\) vs. \(20\)). The direction of Factor B's effect completely reverses depending on the level of Factor A — a classic crossover interaction. Choices A, B, and D all incorrectly claim main effects where marginal means are identical.

Q91. Which of the following best describes a simple random sample (SRS) of size \(n\) from a population?
A Every possible sample of size \(n\) has an equal probability of being selected
B Every individual has an equal probability of being selected, but samples of size \(n\) may not be equally likely
C The population is divided into groups and \(n\) individuals are selected from each group
D Individuals are selected at fixed intervals after a random starting point

An SRS requires that every possible sample of size \(n\) — not just every individual — has an equal chance of selection. Choice B describes a weaker condition that could be satisfied by other methods (e.g., systematic sampling) that do not give all samples equal probability. Choices C and D describe stratified and systematic sampling, respectively.

Q92. A news website posts an online poll asking visitors to click 'yes' or 'no' on a controversial political question, and reports the results as representative of public opinion. Which sampling method does this illustrate?
A Systematic random sampling
B Voluntary response sampling
C Stratified random sampling
D Cluster sampling

Voluntary response sampling occurs when individuals choose whether to participate, typically attracting people with strong opinions. This makes the sample unrepresentative of the broader population. The other choices all involve some form of active random selection by the researcher rather than self-selection by respondents.

Q93. In a clinical trial, some participants receive an inert pill that looks identical to the actual medication being tested. This inert pill is called a:
A Control variable
B Block
C Placebo
D Replication unit

A placebo is a treatment with no active ingredient, designed to look identical to the real treatment. It allows researchers to separate the actual effect of the medication from the psychological effect of believing one is being treated (the placebo effect). A control variable is held constant, and a block is a group of similar experimental units — neither is an inert treatment.

Q94. A school administrator selects every 15th student from an alphabetical enrollment list after randomly choosing a starting position between 1 and 15. This is an example of:
A Cluster sampling
B Stratified random sampling
C Systematic random sampling
D Convenience sampling

Systematic random sampling involves selecting every \(k\)th individual from an ordered list after a random start. Here \(k = 15\). Cluster sampling randomly selects entire groups. Stratified sampling divides the population into subgroups and samples from each. Convenience sampling requires no structured selection process.

Q95. A researcher divides the student population into four class levels (freshman, sophomore, junior, senior) and then randomly selects 50 students from each level. This sampling method is called:
A Cluster sampling
B Multistage sampling
C Systematic sampling
D Stratified random sampling

Stratified random sampling divides the population into non-overlapping subgroups (strata) based on a relevant characteristic, then draws a random sample from each stratum. This differs from cluster sampling, where entire groups are randomly selected rather than individuals from within each group.

Q96. What is the primary purpose of including a control group in a randomized experiment?
A To ensure that each participant has an equal chance of receiving every treatment
B To provide a baseline against which the effect of the treatment can be compared
C To prevent the placebo effect from influencing any participant
D To increase the total number of observations in the study

A control group receives either no treatment or a standard treatment, allowing researchers to attribute any difference in outcomes to the experimental treatment rather than to other factors. Without a baseline for comparison, it is impossible to determine whether observed changes are due to the treatment. The control group does not eliminate the placebo effect — that is addressed through blinding.

Q97. A marketing researcher stands at a busy mall entrance and surveys the first 40 shoppers who agree to participate. This is an example of:
A Stratified random sampling
B Systematic random sampling
C Convenience sampling
D Cluster sampling

Convenience sampling selects individuals who are easiest to reach rather than using a random mechanism. Shoppers at a mall entrance are not representative of the general population — they are available and willing, which introduces selection bias. The other methods all involve formal random selection procedures.

Q98. A variable that is associated with both the explanatory variable and the response variable, and was not accounted for in the study design, is called a:
A Response variable
B Confounding variable
C Blocking variable
D Lurking variable

A confounding variable is one that is related to both the explanatory and response variables and therefore provides an alternative explanation for an observed association. 'Lurking variable' is sometimes used synonymously, but in AP Statistics, a confounding variable specifically refers to one whose effect cannot be separated from the explanatory variable's effect. A blocking variable is intentionally accounted for in the design.

Q99. A polling organization conducts a telephone survey using only landline numbers to estimate the percentage of adults who support a proposed tax increase. Which source of bias is most likely to affect the results?
A Response bias, because respondents may not answer honestly about tax preferences
B Undercoverage bias, because adults who use only mobile phones are excluded from the sampling frame
C Nonresponse bias, because landline owners are less likely to answer unknown calls
D Voluntary response bias, because participants must call in to be included

Undercoverage occurs when parts of the population are systematically excluded from the sampling frame. Adults who rely solely on mobile phones — a group that tends to skew younger — are completely absent from a landline-only frame, making the sample unrepresentative. While nonresponse is also a real concern with telephone surveys (choice C), the most fundamental structural flaw is undercoverage of an entire segment of the population.

Q100. A survey question reads: 'Do you support the reckless plan to cut funding for our schools?' A researcher uses responses to this question to estimate public support for an education budget proposal. Which problem most directly affects the validity of the results?
A Undercoverage, because some households may not be included in the sampling frame
B Nonresponse bias, because people who oppose the plan are less likely to respond
C Response bias due to leading question wording that influences participants toward a particular answer
D Voluntary response bias, because participants chose whether to complete the survey

The phrase 'reckless plan' is loaded language that steers respondents toward opposition, making this a leading question. This creates response bias — a systematic tendency for responses to differ from the true opinion due to the question's wording. Undercoverage and nonresponse are related to who participates, not to how the question itself influences answers.

Q101. A researcher randomly selects 30 school districts from a state and then surveys all teachers within those selected districts. This design is an example of:
A Stratified random sampling
B Simple random sampling
C Cluster sampling
D Systematic sampling

Cluster sampling involves randomly selecting entire groups (clusters) and then including all or most members of those groups. Here the clusters are school districts. This differs from stratified sampling, where the population is divided into subgroups and a random sample is drawn from within each subgroup — in cluster sampling, only selected clusters are sampled, not all of them.

Q102. A researcher observes that students who participate in after-school sports programs have higher GPAs than those who do not. The researcher concludes that sports participation causes higher academic achievement. Which of the following most undermines this conclusion?
A The sample of students may be too small to draw reliable conclusions
B Students who participate in sports may have higher motivation or parental involvement, which could independently raise GPA
C GPA is not a valid measure of academic achievement
D The study should have surveyed teachers rather than examining student records

This is an observational study — students were not randomly assigned to sports participation. Students who choose to participate in sports may differ systematically from those who do not in ways that also affect GPA (e.g., motivation, family support, socioeconomic status). These confounding variables prevent a causal conclusion. Only a randomized experiment can control for such differences.

Q103. In a study testing a new pain reliever, neither the patients nor the physicians measuring pain levels know which patients received the drug and which received a placebo. This experimental design is described as:
A Single-blind, because patients do not know their treatment
B Double-blind, because both patients and evaluators are unaware of treatment assignment
C Placebo-controlled, which automatically eliminates all bias
D Matched pairs, because patients are compared to their own baseline

A double-blind design conceals treatment assignment from both participants and those evaluating outcomes. This prevents the placebo effect from influencing patient-reported pain and prevents evaluator bias from affecting measurements. A single-blind design blinds only one party (usually the participant). The term 'placebo-controlled' refers to the presence of a placebo, not to blinding.

Q104. Stratified random sampling is generally preferred over simple random sampling when:
A The population is too large to create a complete sampling frame
B The researcher wants every possible sample of size \(n\) to be equally likely
C The population contains identifiable subgroups that differ meaningfully on the variable being studied
D The study requires selecting every \(k\)th individual from an ordered list

Stratified sampling improves precision by ensuring representation from each subgroup (stratum). If the strata have different distributions of the response variable, sampling from each stratum separately reduces the variability of estimates compared to an SRS of the same total size. Choice B is a property of SRS, not stratified sampling. Choice D describes systematic sampling.

Q105. In a study examining whether a new fertilizer increases crop yield, researchers pair plots of land with similar soil quality and sun exposure, then randomly assign one plot in each pair to receive the new fertilizer and the other to receive the standard fertilizer. This design is called:
A Completely randomized design
B Matched pairs design
C Cluster randomized design
D Stratified experimental design

A matched pairs design groups experimental units into pairs that are similar on variables that could affect the response (here, soil quality and sun exposure), then randomly assigns one unit in each pair to each treatment. This reduces variability and makes it easier to detect a true treatment effect. A completely randomized design assigns all units to treatments independently without pairing.

Q106. Researchers enroll 1,000 healthy adults in a study and follow them for 10 years, recording their dietary habits and tracking who develops type 2 diabetes. This is an example of:
A A retrospective observational study
B A randomized controlled experiment
C A prospective observational study
D A cross-sectional observational study

A prospective study recruits participants before outcomes occur and follows them forward in time to observe who develops the condition of interest. A retrospective study looks backward at existing records. A cross-sectional study measures exposure and outcome at the same point in time. Since participants are randomly selected but not randomly assigned to diets, this is observational, not experimental.

Q107. A large survey finds that people who carry reusable water bottles report drinking more water per day than those who do not. A journalist concludes that carrying a reusable bottle causes increased water consumption. Which of the following best identifies the flaw in this conclusion?
A The sample size may not be large enough to support a causal claim
B A lurking variable — such as general health consciousness — may explain both the bottle-carrying behavior and higher water intake
C Self-reported water consumption is always unreliable and should not be used in research
D The study should have measured a different response variable to establish causation

Causation cannot be inferred from observational data because lurking variables may be responsible for the observed association. People who carry reusable bottles may generally be more health-conscious, leading them both to use bottles and to drink more water — the bottle itself may not be the cause. Only a randomized experiment (randomly assigning some people to use reusable bottles) could support a causal claim.

Q108. An agricultural researcher wants to test two irrigation methods (drip vs. sprinkler) on tomato yield. The researcher knows that field elevation (low vs. high) significantly affects yield. Which experimental design best controls for elevation?
A Randomly assign all plots to one of the two irrigation methods, ignoring elevation
B Use only low-elevation plots to eliminate elevation as a variable
C Block by elevation and randomly assign irrigation methods within each elevation block
D Pair each low-elevation plot with a high-elevation plot and assign both the same irrigation method

Blocking controls for a known source of variability by grouping experimental units that are similar (same elevation) and then randomly assigning treatments within each block. This allows the researcher to separately estimate the effects of elevation and irrigation method, reducing confounding. Choice A (completely randomized design) does not control for elevation. Choice B reduces external validity by limiting scope. Choice D does not randomly assign different treatments within pairs.

Q109. A researcher examines hospital records of 200 patients who suffered heart attacks and 200 patients who did not, then asks both groups to recall their dietary habits over the past decade. Which methodological concern is most serious in this retrospective study?
A The equal sample sizes in each group make the study results unreliable
B Recall bias may cause heart attack patients to differently reconstruct their dietary history compared to healthy patients
C Hospital records cannot be used as a valid source of patient information
D The study is invalid because it does not include a randomized control group

Recall bias is a major threat in retrospective studies: individuals who experienced a serious health event (heart attack) may be more motivated to search for reasons and therefore recall or report their past diet differently than healthy controls. This systematically distorts the association between diet and heart attack risk. While randomized experiments are stronger for causal inference, observational studies are not invalid — they simply require careful interpretation.

Q110. An experiment randomly assigns 80 volunteers to either receive a new study-skills workshop or to serve as a control. After 8 weeks, the workshop group shows higher exam scores. A critic argues that volunteers who signed up may be more motivated than typical students. This critique most directly challenges which aspect of the study?
A Internal validity, because motivation confounds the effect of the workshop within the study
B External validity, because the findings may not generalize to students who would not voluntarily seek such a workshop
C Reliability, because the study cannot be replicated with different volunteers
D Statistical significance, because motivated students inflate the \(p\)-value

The critic is questioning whether the results can be generalized to the broader student population — this is a question of external validity. Because only motivated volunteers participated, the effect seen may not apply to typical students who would not seek out a study-skills workshop. Internal validity (whether the workshop caused the improvement within this sample) is actually well-supported by random assignment, which controls for differences between groups.

Q111. A researcher uses systematic sampling to select every 20th customer from a database of 4,000 customers sorted by the date of their last purchase, yielding a sample of 200. The database is structured so that purchases tend to peak on the first and third weeks of each month. Which concern about this design is most valid?
A A sample of 200 is too small to draw conclusions from a database of 4,000
B Systematic sampling does not produce a valid probability sample
C If the sampling interval aligns with the periodic purchase cycle, the sample may systematically over- or under-represent customers who buy at peak times
D The random starting point makes this equivalent to a convenience sample

Systematic sampling can introduce periodicity bias when the sampling interval (\(k = 20\)) happens to coincide with a natural cycle in the ordered list. If every 20th record falls disproportionately in peak or off-peak purchase periods, the sample will not accurately represent all customer types. This is a known limitation of systematic sampling and does not affect SRS, which has no fixed interval.

Q112. In a study testing whether a new antidepressant reduces symptoms, patients are randomly assigned to the drug or placebo. After 8 weeks, trained clinicians who know which patients received the drug rate each patient's improvement on a standardized scale. Which flaw most threatens the validity of the clinician ratings?
A The absence of a control group prevents meaningful comparison of outcomes
B Random assignment does not guarantee equal group sizes, distorting the ratings
C Clinician knowledge of treatment assignment may cause them to unconsciously rate drug patients more favorably, introducing measurement bias
D An 8-week study period is universally considered too short to detect antidepressant effects

When evaluators know which treatment a participant received, their assessments can be unconsciously influenced — a form of observer or measurement bias. This is why double-blinding (concealing treatment from both participants and evaluators) is the gold standard in clinical trials. The study is single-blind at best: patients may not know their treatment, but the clinicians do. Random assignment (choice B) addresses group equivalence, not evaluator bias.

Q113. A public health researcher studies the relationship between neighborhood walkability and obesity rates across 50 cities. Cities with higher walkability scores have significantly lower obesity rates. The researcher wants to conclude that improving walkability reduces obesity. Which of the following most severely limits this causal claim?
A Fifty cities is too small a sample for city-level analysis
B Walkability scores are subjective and cannot be used in quantitative analysis
C Wealthier cities tend to have both higher walkability infrastructure and residents with better access to healthcare and nutrition, confounding the association
D Cross-sectional data should never be used to study health outcomes

This is an observational study susceptible to confounding. Wealth is a lurking variable: wealthier cities invest in walkable infrastructure and also have residents with higher incomes, better diets, and greater healthcare access — all of which independently reduce obesity. Without randomly assigning walkability to cities (which is impossible), it is impossible to disentangle walkability's effect from the effects of these confounders.

Q114. A school district implements a new science curriculum in all of its elementary schools simultaneously during the fall semester. Student science test scores improve significantly compared to the previous year. Which alternative explanation (confound) most threatens the conclusion that the new curriculum caused the improvement?
A Students in the district may have had lower test scores the previous year due to random variation or unusual circumstances, making a rebound likely regardless of the curriculum change
B The new curriculum was not implemented in a randomized order across schools
C Science test scores are not a valid measure of curriculum quality
D The district should have surveyed teachers about the curriculum before measuring student outcomes

This is an example of regression to the mean as a confound: if last year's scores were unusually low (perhaps due to disruptions, illness, or other temporary factors), scores would likely rise the following year regardless of any intervention. Without a control group that did not receive the new curriculum, it is impossible to separate the effect of the curriculum from natural score recovery. This is a classic threat to the internal validity of before-after studies.

Q115. Three communities volunteer to participate in a trial of a new public health campaign to reduce smoking rates. Three demographically similar communities serve as controls. After one year, the campaign communities show a larger decrease in smoking rates. A researcher concludes the campaign was effective. Which of the following best explains why this conclusion may not be warranted?
A Three communities per group is sufficient for a valid randomized experiment
B Communities that volunteer to adopt a health campaign may already have greater community motivation to reduce smoking, making their outcomes unrepresentative of typical communities
C Smoking rates are too difficult to measure accurately at the community level
D The study should have used individual-level rather than community-level randomization

Because communities volunteered rather than being randomly assigned, there is self-selection bias: communities willing to adopt a public health campaign may already have populations more motivated or prepared to change behavior. This pre-existing difference — not the campaign itself — could explain the greater smoking reduction. Randomization at the community level (choice D) would help, but only if assignment were truly random rather than based on willingness to participate.

Q116. Which of the following best describes the defining property of a simple random sample (SRS)?
A Every possible sample of size n has an equal probability of being selected
B Each individual in the population has an equal probability of being selected, but some samples are more likely than others
C Members are selected from the population at fixed intervals after a random starting point
D The population is divided into subgroups, and members are randomly chosen from each subgroup

A simple random sample requires that every possible sample of size n has an equal probability of being selected — not merely that each individual has an equal chance. Choice B describes a weaker condition satisfied by other methods such as systematic sampling, where each individual is equally likely to be chosen but not every sample of size n is equally likely. Choice C describes systematic random sampling, and Choice D describes stratified random sampling.

Q117. A television station asks viewers to text in their opinion on whether the city should ban single-use plastics. Which type of bias most threatens the validity of results from this poll?
A Undercoverage bias
B Voluntary response bias
C Nonresponse bias
D Response bias due to question wording

Voluntary response bias occurs when individuals self-select into a survey, rather than being randomly chosen. People who feel strongly — particularly those who oppose the ban — are far more likely to take action and text in. This produces a sample that overrepresents those with strong opinions. Nonresponse bias (Choice C) differs: it occurs when randomly selected individuals fail to respond. Here, no one was randomly selected at all — anyone may participate.

Q118. In a randomized controlled experiment, the primary purpose of including a control group is to:
A Ensure that all participants receive the same treatment so outcomes are comparable
B Provide a baseline against which the effect of the treatment can be compared
C Eliminate the placebo effect so it does not influence the results
D Guarantee that the experiment is conducted under double-blind conditions

A control group receives no active treatment (or receives a placebo) and serves as a baseline comparison. Without it, there is no way to know whether changes in the treatment group result from the treatment itself or from other factors such as the passage of time, the Hawthorne effect, or natural recovery. The control group does not eliminate the placebo effect — it accounts for it by allowing researchers to measure how much improvement occurs even without the active treatment.

Q119. In a single-blind clinical trial, which group is unaware of whether they are receiving the active treatment or the placebo?
A The researchers who administer and evaluate the treatments
B The participants only
C Both the participants and the researchers administering the treatments
D The institutional review board overseeing the study

In a single-blind experiment, the participants do not know which treatment they are receiving. This prevents them from changing their reported outcomes or behavior based on expectations. When both the participants and the researchers interacting with them are unaware of treatment assignment, the design is called double-blind. Double-blinding provides stronger protection against bias because it also prevents researchers from unconsciously treating participants differently or interpreting outcomes in a biased way.

Q120. A journalist wants to estimate the proportion of residents in a large city who support a new transit plan. She interviews the first 50 people she encounters outside her office building. This sampling method is best described as:
A A stratified random sample
B A systematic random sample
C A convenience sample
D A cluster sample

A convenience sample consists of individuals who are easily accessible to the researcher rather than randomly selected from the population. People near the journalist's office building are likely not representative of the entire city — they may share similar occupations, commuting patterns, and neighborhoods. This makes convenience samples prone to bias. A cluster sample (Choice D) would involve randomly selecting groups and surveying all their members, which is not what occurred here.

Q121. In a study examining the relationship between coffee consumption and heart disease, a confounding variable is best described as:
A The amount of coffee consumed daily, because it is the variable being manipulated
B The incidence of heart disease, because it is the outcome being measured
C A variable such as smoking that is associated with both coffee consumption and heart disease risk
D The random assignment of participants to high or low coffee intake groups

A confounding variable is one that is associated with both the explanatory variable (coffee consumption) and the response variable (heart disease), making it difficult to determine whether the explanatory variable itself causes the response. For example, smokers tend to drink more coffee and also have higher heart disease risk — so smoking could create an apparent relationship between coffee and heart disease even if no true causal link exists. Choice A describes an explanatory variable and Choice B describes the response variable.

Q122. The National Agricultural Statistics Service contacts every farm in the United States to collect data on crop yields. This data collection effort is called:
A A stratified random sample, because farms are grouped by crop type
B A census, because data is collected from every member of the population
C A cluster sample, because farms are grouped by geographic region
D A systematic random sample, because farms are surveyed on a regular schedule

A census collects data from every individual in the entire population of interest — here, every farm in the United States. While a census provides complete and accurate information, it is typically expensive, time-consuming, and sometimes logistically impossible for large populations, which is why most large-scale studies use probability sampling methods instead. None of the other choices apply because this is not a sampling procedure at all.

Q123. A researcher studying reading achievement among elementary school students randomly selects 15 schools from a large school district and then collects reading scores from every student in each selected school. This sampling method is best described as:
A Stratified random sampling, because schools serve as natural strata based on grade level
B Simple random sampling, because schools are randomly chosen
C Cluster sampling, because entire randomly selected groups are fully surveyed
D Systematic random sampling, because students are selected in a structured pattern across schools

In cluster sampling, the population is divided into groups called clusters. A random sample of clusters is selected, and all members within the selected clusters are measured. Here, schools are the clusters. This differs critically from stratified sampling, where individuals are randomly selected from within each group (stratum) — not all members of selected groups. Cluster sampling is practical and cost-effective when a population is naturally divided into accessible groups, but it can produce higher variability if clusters differ substantially from one another.

Q124. A national health survey uses random digit dialing to contact households with landline telephones. Which concern is most serious regarding the representativeness of the sample?
A Nonresponse bias, because some people will not answer the phone when called
B Response bias, because questions about health may cause embarrassment and inaccurate answers
C Undercoverage bias, because households relying solely on cell phones are excluded from the sampling frame
D Voluntary response bias, because people choose whether to answer health questions honestly

Undercoverage bias occurs when part of the population has no chance of being selected because they are absent from the sampling frame. Using only landline numbers excludes the roughly 50 to 60 percent of US households that rely solely on cell phones — a group that skews younger, lower income, and may differ in health behaviors. This is a fundamental limitation of the sampling frame itself. Nonresponse bias (Choice A) is also a concern but addresses what happens after someone is contacted, not who can be reached at all.

Q125. In a clinical trial for a new anxiety medication, participants in the placebo group showed a statistically significant reduction in self-reported anxiety scores over 12 weeks. Which of the following best explains this outcome?
A Confounding, because the researchers inadvertently gave some placebo participants the active medication
B The placebo effect, because participants experienced real psychological improvement from believing they were receiving treatment
C Double-blinding, because neither participants nor researchers knew who received the placebo
D Response bias, because participants wanted to please the researchers and reported lower anxiety regardless of actual change

The placebo effect occurs when participants experience genuine, measurable improvements in their condition simply because they believe they are receiving an active treatment. This is a real psychological and sometimes physiological phenomenon, not merely biased reporting. Including a placebo control group is essential in clinical trials precisely to quantify this effect and separate it from the true impact of the medication. Double-blinding (Choice C) is a design feature used to reduce bias, not an explanation for why the placebo group improved.

Q126. A survey asks: 'Do you support wasteful government spending on foreign aid when Americans are struggling at home?' A different survey asks: 'Do you support US foreign aid programs that promote global stability and reduce poverty?' Both surveys address the same underlying policy. The likely difference in responses between the two surveys is best attributed to:
A Undercoverage bias, because different populations received each survey
B Nonresponse bias, because respondents who oppose foreign aid are more likely to respond to one version
C Response bias due to question wording, because emotionally charged or leading language influences respondents' answers
D Sampling variability, because different random samples naturally produce different estimates

Response bias due to question wording occurs when the phrasing of a question systematically leads respondents toward a particular answer. The first version uses loaded terms like 'wasteful' and frames the policy as harming Americans, pushing respondents to oppose it. The second version uses neutral and positive framing. Good survey design requires neutral, balanced language that does not prime respondents toward any answer. This is distinct from sampling variability (Choice D), which refers to random fluctuation between samples, not systematic distortion caused by wording.

Q127. A researcher tracks 800 adults for 10 years and finds that those who sleep fewer than six hours per night have significantly higher rates of type 2 diabetes. The researcher concludes that insufficient sleep causes type 2 diabetes. Which of the following most clearly identifies why this conclusion is not justified?
A The study is too short to observe the full development of type 2 diabetes
B The study is observational, so confounding variables such as diet, exercise, and weight could explain the association without sleep being the cause
C A sample of 800 is insufficient to detect a relationship between sleep and diabetes at the population level
D The researcher should have measured insulin resistance rather than diabetes diagnosis to establish causation

In an observational study, the researcher does not control or randomly assign the explanatory variable (sleep duration). People who sleep fewer than six hours may differ in many other ways — they may have more stressful jobs, poorer diets, less exercise, or higher BMI — all of which are associated with diabetes risk. These confounders could fully or partially explain the observed association. Only a randomized experiment, where sleep duration is experimentally assigned and other variables are controlled, would support a causal conclusion. Choice C is incorrect because 800 is a reasonable sample size for detecting such associations.

Q128. An agricultural researcher tests a new drought-resistant wheat variety by planting it in a single one-acre plot and comparing the yield to a single adjacent plot of standard wheat. A statistician suggests the experiment lacks replication. What does this mean?
A The experiment should be repeated by an independent research team before results are trusted
B Each treatment should be applied to multiple experimental units so that results reflect the treatment effect rather than the unique conditions of one plot
C Both wheat varieties should be planted in the same plot to control for soil differences
D The researcher needs a control group that receives no wheat planting at all

Replication in experimental design means applying each treatment to multiple experimental units (here, multiple plots). With only one plot per treatment, any difference in yield could easily be due to random variation in that particular plot — its specific soil conditions, drainage, or sun exposure — rather than the wheat variety itself. Using multiple plots per treatment allows researchers to estimate the natural variability in yield and determine whether the observed difference is larger than what chance alone would produce. Choice A describes scientific reproducibility, which is a different concept.

Q129. A researcher studying the effect of caffeine on short-term memory recruits 40 pairs of identical twins. Within each pair, one twin is randomly assigned to consume caffeine before a memory test while the other consumes a decaffeinated beverage. The primary advantage of this matched pairs design over a completely randomized design is that it:
A Eliminates the placebo effect because twins are genetically identical and respond identically to placebos
B Controls for genetic variation, so differences in memory scores within pairs are more likely attributable to caffeine
C Allows each participant to serve as their own control, doubling the effective sample size
D Removes the need to randomly assign treatments because twins are already matched

Identical twins share the same DNA, so pairing them controls for genetic differences that might otherwise affect memory performance. By randomly assigning one twin in each pair to caffeine and the other to decaf, the researcher ensures that any systematic difference in memory scores within pairs is more likely due to caffeine rather than underlying genetic variation. A completely randomized design could by chance place twins with naturally stronger memories disproportionately in one treatment group. Random assignment is still used (Choice D is incorrect) — it occurs within each pair, not between pairs.

Q130. A political polling organization mails surveys to 5,000 randomly selected registered voters asking about their views on a proposed tax increase. Only 800 surveys are returned. Which concern is most directly related to nonresponse bias?
A The 800 respondents may differ systematically from the 4,200 non-respondents in their views on taxation
B A response rate of 16% is below the standard threshold for a valid survey and automatically invalidates the results
C Using only registered voters excludes unregistered eligible voters who may have different tax preferences
D A sample of 800 is too small to represent the views of all registered voters in a national election

Nonresponse bias occurs when individuals who respond to a survey differ systematically from those who do not. Voters who feel strongly about a tax increase — particularly those who oppose it — may be more motivated to return the survey, while indifferent voters may discard it. This creates a sample skewed toward strong opinions. There is no universal response rate threshold that invalidates results (Choice B is incorrect); what matters is whether respondents are representative. Choice C describes undercoverage, a different bias introduced by the sampling frame, not by nonresponse.

Q131. A consumer research firm finds that satisfaction ratings for a new product are significantly higher when participants are asked about overall satisfaction before being asked about specific complaints, compared to when the complaint questions come first. This difference in results is best explained by:
A Voluntary response bias, because satisfied customers are more likely to participate in both question orders
B Undercoverage bias, because dissatisfied customers are not included in either version of the survey
C Response bias due to question order effects, because the context established by earlier questions influences responses to later ones
D Sampling variability, because different groups of randomly selected participants will naturally produce different satisfaction scores

Question order effects are a form of response bias in which earlier questions prime or anchor respondents, influencing how they interpret and answer subsequent questions. When specific complaints are raised first, respondents are primed to think critically, which lowers overall satisfaction ratings. When overall satisfaction is asked first, respondents give an impressionistic positive rating before complaints are mentioned. This is why well-designed surveys carefully consider question ordering, sometimes randomizing or counterbalancing the sequence. Sampling variability (Choice D) refers to random fluctuation between samples, not systematic differences caused by design.

Q132. A researcher testing a new blood pressure medication divides participants into two blocks — those under 50 years old and those 50 and older — before randomly assigning participants within each block to treatment or control. The primary reason for blocking on age is to:
A Ensure that younger and older participants receive different doses of the medication
B Prevent older participants from being assigned to the control group
C Reduce variability in treatment comparisons by grouping participants who are similar with respect to a variable related to blood pressure
D Allow the researcher to draw separate conclusions about the drug's effectiveness for younger and older adults

Blocking groups experimental units that are similar with respect to a variable that is likely related to the response (here, age affects baseline blood pressure and cardiovascular health). By assigning treatments randomly within each block, the comparison between treatment and control occurs among people similar in age. This reduces the variability that would otherwise make it harder to detect the true treatment effect. Drawing separate conclusions for each age group (Choice D) is a secondary possibility, but improving precision across the entire study is the primary purpose of blocking.

Q133. A researcher wants to estimate average weekly screen time for middle school students in a large district. She randomly selects 30 students from each of the district's 12 schools and surveys them. Her colleague suggests instead randomly selecting 4 schools and surveying every student in those schools. Which statement correctly compares these two designs?
A The colleague's approach is stratified sampling; the researcher's approach is cluster sampling
B The researcher's approach is stratified sampling; the colleague's approach is cluster sampling
C Both approaches are equivalent forms of probability sampling and will produce equally precise estimates
D The researcher's approach guarantees a more biased estimate because selecting from each school overrepresents large schools

The researcher's design is stratified random sampling: the population is divided into strata (schools), and a random sample is drawn from each stratum, ensuring all schools are represented. The colleague's design is cluster sampling: entire groups (schools) are randomly selected and all members surveyed. Stratified sampling typically produces more precise estimates when the strata are internally similar with respect to the variable of interest, because it guarantees representation from every group. Cluster sampling is more convenient but can be less precise if clusters differ from each other in screen time habits.

Q134. A research organization uses a list of active social media users as its sampling frame for a survey about smartphone usage habits among US adults. Which limitation of this sampling frame is most serious?
A Social media users may have multiple accounts, causing some individuals to have higher probability of selection
B Adults who do not use social media are excluded from the frame, and this group may systematically differ in smartphone usage patterns from social media users
C Social media platforms collect demographic data that could compromise survey anonymity
D The sampling frame changes constantly as users join or leave platforms, making it impossible to define a fixed population

Undercoverage bias at the sampling frame level occurs when a segment of the target population has no chance of selection. Adults who do not use social media tend to be older, less digitally engaged, and may have very different smartphone usage patterns — possibly far less screen time or fewer apps. Because this group is systematically excluded and differs meaningfully from social media users, estimates of 'US adults' smartphone habits based solely on social media users will be biased. Choice A (multiple accounts) causes unequal probability of selection but does not exclude entire segments, so it is a less fundamental flaw.

Q135. A data analyst observes a strong positive correlation between the number of swimming pool drownings per month and monthly ice cream sales. A colleague claims this shows ice cream makes people reckless near water. A statistician argues a lurking variable explains the association. Which option most accurately identifies the lurking variable and mechanism?
A Geographic region is a lurking variable, because coastal cities have both more pools and more ice cream shops
B Population size is a lurking variable, because larger cities have more of everything including drownings and ice cream sales
C Seasonal temperature is a lurking variable, because warm weather independently increases both swimming activity and ice cream consumption
D Media coverage is a lurking variable, because news reports about drownings cause stress that increases ice cream consumption

Seasonal temperature is the classic lurking variable in this scenario. During warmer months, people both purchase more ice cream and spend more time swimming — directly increasing the opportunity for drowning accidents. Ice cream sales and drowning rates are both responses to hot weather and are not causally related to each other. This illustrates a fundamental principle: correlation does not imply causation. A lurking variable can create a spurious association between two variables that have no direct causal link. Population size (Choice B) might explain city-level differences but does not explain month-to-month variation within a region.

Q136. Researchers recruit 150 college athletes via a campus athletics department email and randomly assign them to either a new protein supplement or a placebo for 10 weeks. The supplement group shows significantly greater muscle mass gains. Which conclusion is best supported by this study design?
A The protein supplement causes greater muscle mass gains in all college students, because random assignment was used
B The protein supplement causes greater muscle mass gains in college athletes who would participate in such a study, but results may not generalize to all college students or non-athletes
C The protein supplement is associated with greater muscle mass gains, but no causal conclusion is possible because participants knew they were in an experiment
D No valid conclusion can be drawn because the study was conducted on a volunteer sample

Random assignment of treatments supports a causal conclusion — it controls for confounding variables, so the observed difference in muscle gains is attributable to the supplement. However, because participants were recruited from a specific population (college athletes who volunteered), the results can only be generalized to similar individuals, not all college students or the general public. This reflects the distinction between internal validity (supported by random assignment) and external validity (limited by the volunteer, athlete-specific sample). Choice C incorrectly denies causation — random assignment is what enables it, regardless of whether participants know they are in a study.

Q137. A gym owner wants to test whether a new high-intensity interval training (HIIT) program improves cardiovascular fitness more than standard aerobic exercise. She assigns members who sign up in January to the HIIT program and members who sign up in February to the standard program, then measures fitness improvements after 8 weeks. Which design flaw most seriously threatens the causal interpretation of the results?
A Eight weeks is too short to produce measurable cardiovascular fitness improvements in either group
B The gym owner should have used a double-blind design to prevent participants from knowing their program type
C January and February sign-ups may differ systematically in motivation, fitness level, or goals, so the groups are not comparable before the programs begin
D The study needs a third group that does no exercise at all to serve as a true control

The critical flaw is that treatment assignment is based on when members signed up, not on random assignment. January sign-ups may be driven by New Year's resolutions and may differ from February sign-ups in baseline motivation, fitness level, and commitment — all of which could affect fitness gains independently of the program. This is a form of selection bias. Without random assignment, the groups are not comparable, so any difference in outcomes cannot be confidently attributed to the HIIT program. Double-blinding (Choice B) would be ideal but is a secondary concern — the non-random assignment is the more fundamental threat.

Q138. A university reports that its overall acceptance rate for female applicants (38%) is lower than for male applicants (45%). However, when acceptance rates are broken down by college within the university, female applicants have equal or higher acceptance rates than male applicants in every individual college. Which of the following best explains this apparent contradiction?
A The university's aggregate data contains a computational error that is corrected by disaggregating by college
B Female applicants tend to apply in greater numbers to the more selective colleges within the university, so the aggregate female acceptance rate is lower even though women are accepted at equal or higher rates within each college
C Male applicants have stronger academic qualifications on average, making department-level comparisons misleading
D Affirmative action policies at the individual college level correct for bias that appears in the aggregate data

This is a classic example of Simpson's paradox, where a trend present in aggregated data reverses when data is broken into subgroups. If female applicants disproportionately apply to highly selective colleges (such as law or medicine) that have low acceptance rates for all applicants, the overall female acceptance rate will be lower — even if women succeed at equal or higher rates within each college. The college applied to acts as a lurking variable associated with both the applicant's gender and the probability of acceptance. Analyzing only aggregate data without accounting for this lurking variable produces a misleading picture of admissions bias.

Q139. An observational study reports that adults who own pets have significantly lower blood pressure than adults without pets and concludes that pet ownership reduces blood pressure. A reviewer argues that physical activity level is a confounding variable. Which explanation best supports the reviewer's position?
A More physically active people are more likely to own pets (particularly dogs requiring walks) and also have lower blood pressure due to exercise, so physical activity could explain the association without pet ownership being the cause
B Pet owners may be more likely to report lower blood pressure to appear healthier, introducing response bias into the study
C People with naturally lower blood pressure may have more energy, making them more likely to adopt pets, reversing the causal direction
D The study should have randomly assigned participants to own or not own pets to eliminate physical activity as a confounder

For a variable to be a confounder, it must be associated with both the explanatory variable (pet ownership) and the response variable (blood pressure). Physical activity meets both criteria: people who are more active are more likely to own pets that require exercise (such as dogs), and exercise independently lowers blood pressure. This means physical activity could fully or partially explain the observed difference in blood pressure between pet owners and non-owners without pet ownership causing any direct benefit. Choice C describes reverse causation — a different validity threat — and Choice D describes how to improve the study design, not how physical activity acts as a confounder in the current study.

Q140. A national polling firm uses a three-stage sampling procedure: randomly selecting 80 counties from all US counties, then randomly selecting 6 census tracts within each selected county, then randomly selecting 25 households within each selected census tract. Compared to a simple random sample of the same total number of households, the most important statistical disadvantage of this multi-stage design is that:
A Multi-stage sampling cannot produce unbiased estimates of national population parameters
B Households within the same census tract tend to be more similar to each other than to households selected nationally at random, reducing the effective sample size and increasing the margin of error
C Randomly selecting counties first means some states will have no representation, creating systematic geographic bias
D The complexity of the design makes it impossible to calculate a valid sampling distribution for the estimator

In multi-stage (cluster-based) sampling, the final sampled units are geographically clustered. Households within the same census tract tend to share similar income levels, racial and ethnic composition, housing type, and political views — they are more alike than a set of households drawn at random from across the country. This within-cluster similarity means each additional household from the same tract adds less new information than an independently selected household would. The result is that the effective sample size is smaller than the nominal sample size, producing wider confidence intervals than a true SRS of the same count. Despite this, multi-stage sampling is far more cost-efficient for national surveys and does produce valid (though less precise) probability-based estimates.

Q141. A researcher divides a school's student population into four grade levels and then randomly selects 50 students from each grade level to survey. Which sampling method is being used?
A Stratified random sampling
B Cluster sampling
C Systematic sampling
D Simple random sampling

Stratified random sampling divides the population into non-overlapping subgroups called strata (here, grade levels) and then takes a random sample from each stratum. Cluster sampling would involve randomly selecting entire grade levels and surveying everyone in them. Simple random sampling would not guarantee representation from each grade level.

Q142. A school principal wants data on every enrolled student's GPA, so she collects GPA information from all 1,200 students currently attending. This data collection method is best described as a:
A Census
B Simple random sample
C Convenience sample
D Cluster sample

A census collects data from every member of the population of interest. Since the principal is measuring all 1,200 enrolled students — not a subset — this is a census. A sample only includes a subset of the population.

Q143. A radio station asks listeners to call in and vote on whether local parking fees should be increased. Which type of bias most likely affects the results of this poll?
A Voluntary response bias
B Undercoverage bias
C Response bias
D Interviewer bias

Voluntary response bias occurs when individuals choose whether to participate, and those who feel most strongly about an issue are more likely to respond. Callers who oppose increased parking fees are disproportionately motivated to call in. This is different from nonresponse bias, which occurs when selected individuals fail to respond to a survey they were asked to complete.

Q144. Which of the following characteristics most clearly distinguishes an experiment from an observational study?
A In an experiment, the researcher deliberately imposes a treatment on subjects
B In an experiment, the researcher only observes and records behavior without interference
C Experiments always involve larger sample sizes than observational studies
D Only experiments can be used to identify associations between variables

The defining feature of an experiment is that the researcher actively assigns treatments to subjects. In an observational study, the researcher simply measures variables as they naturally occur without intervening. Both study types can identify associations, but only well-designed experiments support causal conclusions.

Q145. In a clinical trial testing a new headache medication, one group receives the medication and another group receives a sugar pill that looks identical to the medication. What is the primary purpose of the group receiving the sugar pill?
A To provide a baseline for comparison so that the effect of the medication can be isolated
B To ensure the sample size is large enough for statistical significance
C To test whether sugar pills independently reduce headache symptoms
D To make the experiment double-blind without additional procedures

A control group receiving a placebo provides a baseline response against which the treatment group can be compared. Without a control group, researchers cannot tell how much improvement is due to the drug versus natural recovery, the placebo effect, or other factors. The control group does not by itself create a double-blind design.

Q146. In a drug trial, a group of participants who received an inactive sugar pill reported significant improvement in their symptoms. This phenomenon is called the:
A Placebo effect
B Hawthorne effect
C Confounding effect
D Regression to the mean

The placebo effect occurs when subjects experience real changes in symptoms simply because they believe they are receiving treatment, even when the treatment is inert. The Hawthorne effect refers to behavior changes caused by the awareness of being observed, which is a related but distinct phenomenon.

Q147. A quality control inspector at a factory selects every 20th item coming off the production line for inspection. Which sampling method is this?
A Systematic sampling
B Stratified random sampling
C Cluster sampling
D Convenience sampling

Systematic sampling involves selecting every \(k\)th individual from a list or sequence after a random starting point. Here, every 20th item is selected, making this a systematic sample. Convenience sampling would involve selecting items that are easiest to access, not every \(k\)th item.

Q148. Why is replication considered an important principle of a well-designed experiment?
A It allows researchers to estimate natural variability and increases confidence that observed effects are real
B It ensures that all participants are assigned to the same treatment group for consistency
C It eliminates the need for a control group by repeating the treatment multiple times
D It guarantees that the experiment will produce statistically significant results

Replication means applying each treatment to multiple subjects (or repeating the experiment). This allows researchers to distinguish true treatment effects from random chance variation. Without replication, it is impossible to know whether an observed difference reflects the treatment or just natural individual-to-individual variability. Replication does not eliminate the need for a control group.

Q149. A researcher wants to study the study habits of high school students across a large city. She randomly selects 12 schools from the city and then surveys every student in those selected schools. Which sampling method is she using?
A Cluster sampling
B Stratified random sampling
C Simple random sampling
D Systematic sampling

Cluster sampling divides the population into groups (clusters) and randomly selects entire clusters to study. Here, schools are the clusters, and every student in the selected schools is surveyed. Stratified random sampling would require selecting a random subset of students from every school, not all students from only some schools.

Q150. A university mails a satisfaction survey to 1,000 randomly selected alumni. Only 180 surveys are returned, and those respondents report an average satisfaction score of \(4.2\) out of \(5\). Why might this result be misleading?
A Nonresponse bias may make the sample unrepresentative, since alumni who chose to respond may have systematically different opinions than those who did not
B The sample size of 180 is too small to draw any conclusions regardless of the response rate
C Mail surveys always produce inflated satisfaction ratings compared to online or phone surveys
D The result is valid because the original 1,000 were selected randomly

Nonresponse bias occurs when those who respond differ systematically from those who do not. Alumni who feel strongly (either positively or negatively) may be more motivated to return the survey than those with neutral opinions. Although the original 1,000 were randomly selected, only an \(18\%\) response rate raises serious concerns about whether the 180 respondents represent the full group.

Q151. In a study testing a new anxiety medication, neither the patients nor the physicians evaluating their progress know which patients received the real drug and which received a placebo. This experimental design is best described as:
A Double-blind, because neither the subjects nor the evaluators know the treatment assignments
B Single-blind, because only the patients are unaware of their treatment
C Randomized, because patients were randomly assigned to treatment groups
D Controlled, because a placebo group is included in the study

A double-blind design keeps both the participants and the researchers evaluating outcomes unaware of who received which treatment. This prevents both the placebo effect (from subjects) and evaluator bias (from researchers unconsciously rating treated patients more favorably). A single-blind design would only hide the assignment from participants.

Q152. A researcher tests whether listening to classical music improves focus. Each participant completes a baseline focus assessment, listens to classical music for 20 minutes, and then completes a second equivalent focus assessment. Which experimental design does this represent?
A Matched pairs design, because each subject serves as their own control across the two conditions
B Completely randomized design, because all subjects complete both assessments
C Block design, because subjects are grouped before receiving treatment
D Stratified design, because the same assessment instrument is used for all subjects

In a matched pairs design, each subject is measured under both conditions (before and after), so each person serves as their own control. This controls for individual differences in baseline focus. A completely randomized design would randomly assign different subjects to either the music condition or a no-music condition.

Q153. A study finds that cities with more fast food restaurants have higher rates of cardiovascular disease. A researcher concludes that fast food restaurants cause cardiovascular disease. What is the most important problem with this causal conclusion?
A A lurking variable such as population density or socioeconomic status may be associated with both the number of fast food restaurants and cardiovascular disease rates
B The study should have used a randomized experiment rather than an observational study to establish causation
C Cardiovascular disease rates are too difficult to measure accurately across cities
D The number of fast food restaurants in a city is not a meaningful variable for a health study

A lurking (or confounding) variable is one that is associated with both the explanatory and response variable, creating the appearance of a causal relationship when none may exist. Denser, lower-income urban areas may have both more fast food restaurants and higher cardiovascular disease rates due to diet, access to healthcare, and other factors. While an experiment would be ideal, the primary flaw stated here is the failure to account for confounding.

Q154. A researcher stands outside a gym at 6 AM and surveys the first 50 people who enter in order to study Americans' exercise habits. What is the primary flaw in this approach?
A This is a convenience sample that is unlikely to represent all Americans, since early-morning gym-goers are systematically more fitness-oriented than the general population
B The researcher should have used cluster sampling by randomly selecting gyms nationwide
C A sample of 50 people is always too small to draw meaningful conclusions about exercise habits
D Surveying people as they enter a location always introduces response bias

Convenience sampling selects individuals who are easy to reach, but this produces a sample that may not represent the target population. People who attend a gym at 6 AM likely have stronger exercise habits than average Americans, so conclusions drawn from this group cannot be generalized. The sample size alone does not determine representativeness — a well-designed random sample of 50 could be more useful than this larger convenience sample.

Q155. A researcher tests three study techniques (rereading, practice testing, and concept mapping) on high school students. Because she believes grade level may affect the results, she randomly assigns students within each grade level to one of the three techniques. This is an example of:
A A randomized block design, where grade level is the blocking variable
B A stratified random sample, where grade level is the stratification variable
C A matched pairs design, where students are paired by grade level
D A completely randomized design with three treatment groups

In a randomized block design, subjects are first grouped into blocks based on a variable expected to affect the response (here, grade level), and then randomly assigned to treatments within each block. This reduces variability and allows the treatment effect to be estimated more precisely. Stratified sampling is a technique for data collection, not an experimental design.

Q156. A health department surveys residents about access to fresh produce by calling landline telephone numbers drawn from the local directory. Which form of bias is most likely to affect this survey?
A Undercoverage, because households without landlines — often lower-income or younger residents — are excluded from the sampling frame
B Voluntary response bias, because only people who want to discuss food access will answer the call
C Response bias, because questions about food access elicit socially desirable answers
D Nonresponse bias, because some households that own landlines will not answer the phone

Undercoverage occurs when some members of the population have no chance of being selected because they are not included in the sampling frame. Households without landlines cannot be reached by this method. Since landline ownership correlates with age and income, the sample systematically excludes certain demographic groups. Note that nonresponse bias (choice D) is also possible but requires the household to be in the frame first; undercoverage is the more fundamental issue here.

Q157. Two surveys are administered to equivalent random samples from the same population. Survey 1 asks: 'Given the rise in violent crime, do you support increased funding for local police?' Survey 2 asks: 'Do you support increased funding for local police?' Which outcome is most likely?
A Survey 1 will yield higher support for police funding because the leading premise about rising crime primes respondents toward a particular answer
B Both surveys will yield similar results because underlying opinions about police funding are stable and unaffected by question wording
C Survey 2 will yield higher support because respondents are not primed to think about crime before answering
D Neither survey result is valid because public opinion questions always introduce too much subjectivity

Response bias can result from question wording that leads respondents toward a particular answer. Survey 1 introduces a premise — rising violent crime — that frames the question in a way likely to increase support for police funding. This is called a leading question. Question wording effects are well-documented in survey research and can produce meaningfully different results even from equivalent random samples.

Q158. A hospital system reports that Hospital A has a lower overall patient mortality rate than Hospital B. However, when patients are separated by condition severity (mild versus serious cases), Hospital B has a lower mortality rate in both categories. What statistical phenomenon explains this apparent contradiction?
A Simpson's paradox, where an association seen in combined data reverses when data are broken into subgroups, often because a confounding variable (case mix) influences both group membership and the outcome
B Nonresponse bias, because not all patient deaths are recorded consistently in hospital databases
C Voluntary response bias, because sicker patients self-select into Hospital A
D The ecological fallacy, because individual-level conclusions cannot be drawn from group-level data

Simpson's paradox occurs when a trend present in aggregated data disappears or reverses when the data are disaggregated into subgroups. Here, Hospital A may treat a higher proportion of mild cases (which have lower mortality regardless of hospital quality), making its overall rate look better. Once case severity is controlled for, Hospital B performs better in both categories. The confounding variable is case-mix severity.

Q159. An observational study finds that people who carry lighters are significantly more likely to develop lung cancer than those who do not. A researcher claims that carrying a lighter causes lung cancer. Which of the following best explains the critical flaw in this causal conclusion?
A Smoking is a confounding variable associated with both carrying a lighter and developing lung cancer, making it impossible to attribute lung cancer risk to lighter-carrying itself
B The study should have used stratified random sampling to ensure lighter-carriers and non-carriers were equally represented in the sample
C Lung cancer rates are too variable across populations to make reliable comparisons in observational studies
D The conclusion is valid as long as the sample size was large enough to detect a statistically significant association

A confounding variable is associated with both the explanatory variable (carrying a lighter) and the response variable (lung cancer). People who smoke are much more likely to carry lighters and to develop lung cancer. The apparent relationship between lighters and cancer is not causal — it is explained by the lurking variable of smoking. Statistical significance alone (choice D) does not establish causation, especially in observational studies.

Q160. A school district tests a new math curriculum by implementing it in schools that voluntarily opt in, and compares their students' test scores to schools continuing with the old curriculum. Students in new-curriculum schools score significantly higher. Which is the most serious threat to the validity of this conclusion?
A Self-selection bias: schools that opted in may already have more motivated teachers, stronger administrative support, or higher-achieving students, so the score difference cannot be attributed solely to the curriculum
B The study would have been internally valid if a larger number of schools had been included in each group
C Standardized test scores are not a reliable measure of mathematical understanding and should not be used to evaluate curricula
D The study lacks a placebo condition, since students in both groups know which curriculum they are using

When subjects (or institutions) self-select into treatment conditions, the groups may differ in important ways before the treatment begins. Schools motivated enough to volunteer for a new curriculum may differ systematically from those that did not, making a fair comparison impossible. This is a lack of random assignment — the most fundamental requirement for drawing causal conclusions. Increasing the sample size (choice B) would not resolve this confounding.

Q161. A researcher needs to estimate average household income across a large city whose neighborhoods vary dramatically in wealth. Which sampling strategy would most efficiently produce precise estimates while properly accounting for this variation?
A Stratified random sampling using neighborhoods as strata, because it ensures proportional representation from each wealth level and reduces within-stratum variability relative to simple random sampling
B Cluster sampling using neighborhoods as clusters, because randomly selecting entire neighborhoods is always more efficient than stratified sampling when geographic groupings exist
C Simple random sampling, because giving every household an equal probability of selection eliminates all forms of sampling bias
D Systematic sampling, because selecting every \(k\)th household from a city directory is equally precise to stratified sampling and easier to implement

When a population contains clearly defined subgroups that are internally homogeneous but differ from each other (high-income vs. low-income neighborhoods), stratified sampling reduces sampling variability compared to simple random sampling. By sampling within each stratum, the researcher captures the full range of the variable of interest. Cluster sampling (choice B) often increases variability because individuals within a cluster (neighborhood) tend to be similar to each other.

Q162. A researcher conducts a well-designed randomized controlled experiment on the effects of sleep deprivation on decision-making, using 120 college students recruited from an introductory psychology course. The experiment is internally valid. Which statement most accurately assesses its external validity?
A The results may not generalize to the broader population because psychology course volunteers may differ from the general adult population in age, sleep patterns, and cognitive demands
B Because participants were randomly assigned to conditions, the results are valid for all human populations
C The results are externally valid as long as a sample of 120 is considered statistically sufficient for the analysis
D External validity is not a concern for experiments; it applies only to observational studies

External validity (generalizability) refers to whether findings from a study can be applied to other populations, settings, or times. Internal validity — achieved here through random assignment — only ensures that the observed effect is attributable to the treatment within this study. College students in psychology classes are a convenience sample and may not represent working adults, older populations, or those with different sleep routines. External validity is a concern for all research designs, not just observational studies.

Q163. A researcher compares two teaching methods (lecture and active learning) in two different schools — one urban and one rural. Active learning dramatically outperforms lecture in the urban school, but lecture slightly outperforms active learning in the rural school. What does this pattern most strongly suggest?
A There is an interaction effect between teaching method and school type, meaning the effectiveness of the teaching method depends on which type of school it is used in
B The results are inconclusive because neither method consistently outperforms the other across both schools
C The urban school results are outliers and should be discarded because they deviate too far from the rural school results
D Active learning is definitively the superior method because it produced a larger effect size in the urban school

An interaction effect (also called effect modification) occurs when the effect of one variable depends on the level of another variable. Here, the advantage of active learning depends on school type: large in urban schools, reversed in rural schools. This means the two factors do not operate independently. Concluding that active learning is superior overall (choice D) would ignore this interaction and could lead to harmful policy decisions in rural contexts.

Q164. A polling organization uses this procedure: randomly select 100 counties from all U.S. counties; within each selected county, randomly select 5 ZIP codes; within each ZIP code, randomly select 20 registered voters. Which of the following correctly identifies a key statistical limitation of this multi-stage cluster design compared to a simple random sample of the same total size (\(100 \times 5 \times 20 = 10{,}000\) voters)?
A The multi-stage design tends to produce higher sampling variability than a simple random sample of equal size, because individuals within the same cluster tend to be more similar to one another than to the broader population
B The multi-stage design is statistically invalid because it does not give every registered voter an equal probability of selection
C The multi-stage design produces lower sampling variability than a simple random sample because clustering geographically groups similar respondents, reducing measurement error
D The multi-stage design introduces response bias because voters in the same ZIP code share political views

A fundamental trade-off of cluster sampling is the design effect: because individuals within the same cluster (county or ZIP code) tend to share demographic, economic, and political characteristics, the effective sample size is smaller than the nominal sample size. This inflates sampling variability relative to a true simple random sample of the same \(n\). Despite this limitation, multi-stage designs are used in practice because they are far less costly than a true national simple random sample.

Q165. A randomized, double-blind, placebo-controlled trial of a new antidepressant shows that \(35\%\) of the drug group and \(28\%\) of the placebo group improved at 8 weeks. However, \(40\%\) of all enrolled participants dropped out before the 8-week endpoint, and dropout was more common in the drug group due to side effects. The researchers analyze only participants who completed the study. Which of the following best identifies the primary threat to the validity of the reported conclusion?
A Attrition bias: participants who dropped out due to side effects are systematically different from completers, and excluding them likely overstates the drug's effectiveness
B Placebo effect: the \(28\%\) improvement in the placebo group is too large to be explained by chance and suggests the drug has no real effect
C Lack of blinding: participants who experienced side effects likely deduced they were in the drug group, invalidating the double-blind design
D Undercoverage: the dropouts were excluded from the original sampling frame before the trial began

Attrition bias (differential dropout) is a serious threat to internal validity in longitudinal studies. When dropout is related to treatment assignment and outcome — here, sicker or less tolerant participants in the drug group left early — analyzing only completers creates groups that are no longer comparable. The drug group completers may be those who tolerated the drug best, making the \(35\%\) improvement rate an overestimate of effectiveness in the full enrolled population. A complete-case analysis violates the intent-to-treat principle. Choice C raises a valid concern about unblinding, but the more direct and primary threat given the \(40\%\) differential dropout rate is attrition bias.

Q166. What is the key feature that distinguishes an experiment from an observational study?
A Experiments always use larger sample sizes than observational studies
B Experiments randomly assign subjects to treatment conditions
C Experiments measure more variables than observational studies
D Experiments are always conducted in a laboratory setting

The defining feature of an experiment is that the researcher actively imposes treatments by randomly assigning subjects to conditions. This random assignment allows researchers to make causal conclusions. Observational studies simply observe subjects without intervening, so any association found cannot be attributed to cause and effect. Sample size, number of variables, and setting are not what define the difference.

Q167. In a simple random sample (SRS) of size \(n\) drawn from a population of size \(N\), which of the following properties must hold?
A Every individual has an equal probability \(\frac{1}{N}\) of being selected
B Every possible sample of size \(n\) has an equal probability of being chosen
C Individuals are selected at fixed intervals from an ordered list
D The population is divided into groups before individuals are chosen

An SRS requires that every possible sample of size \(n\) has an equal probability of being selected — this is a stronger condition than simply saying each individual has equal probability. Choice A describes equal inclusion probability, which is a consequence of an SRS but does not fully define it. Systematic sampling (choice C) and stratified sampling (choice D) are distinct methods that do not necessarily give every sample of size \(n\) equal probability.

Q168. A researcher divides a city into 20 neighborhoods and randomly selects 4 neighborhoods. Every resident in those 4 neighborhoods is then surveyed. Which sampling method is this?
A Stratified random sampling
B Systematic random sampling
C Cluster sampling
D Convenience sampling

Cluster sampling divides the population into groups (clusters), randomly selects some clusters, and then includes all members of the selected clusters. Here, neighborhoods are the clusters. In stratified sampling (choice A), you sample from every group, not just selected ones. Systematic sampling (choice B) selects every \(k\)th individual from a list. Convenience sampling (choice D) uses whoever is easiest to reach, without a random mechanism.

Q169. A survey is mailed to 2,000 randomly selected households, but only 400 are returned. The non-returned surveys most directly introduce which type of bias?
A Response bias
B Undercoverage bias
C Nonresponse bias
D Voluntary response bias

Nonresponse bias occurs when individuals who do not respond differ systematically from those who do. Here, the 1,600 households that did not return surveys may have different opinions or characteristics than the 400 who responded. Response bias (choice A) refers to inaccurate answers from those who do respond. Undercoverage (choice B) means some members of the population had no chance of being included. Voluntary response (choice D) occurs when people choose themselves, rather than being chosen.

Q170. In a well-designed experiment, the group that does not receive the active treatment but is otherwise treated identically to the treatment group is called the:
A Placebo group
B Control group
C Block group
D Replication group

The control group serves as a baseline for comparison and does not receive the active treatment. It may or may not receive a placebo. The placebo group (choice A) is a specific type of control group that receives an inactive treatment, making it a subset of the broader concept of a control group. Blocking (choice C) is a design technique to reduce variability. Replication (choice D) refers to repeating the experiment or applying each treatment to multiple subjects.

Q171. A local radio station asks listeners to call in and vote on whether the mayor is doing a good job. Which type of sample does this represent?
A Stratified random sample
B Systematic random sample
C Voluntary response sample
D Cluster sample

A voluntary response sample consists of people who choose to participate on their own. Such samples are biased because people with strong opinions — especially negative ones — are more likely to respond, making the results unrepresentative. None of the random sampling methods (choices A, B, or D) involve self-selection; they all use a defined random mechanism to determine who is included.

Q172. In a study, a variable that is associated with both the explanatory variable and the response variable, and therefore can create a misleading appearance of causation, is called a:
A Confounding variable
B Response variable
C Blocking variable
D Placebo variable

A confounding variable (also called a lurking variable) is related to both the explanatory and response variables, making it difficult to determine whether the explanatory variable truly causes changes in the response. For example, if wealthier people both exercise more and eat better, wealth confounds a study of diet and health. Blocking variables (choice C) are used in experimental design to control for known sources of variability, not to describe bias in observational studies.

Q173. A researcher wants to compare the reading levels of students across four grade levels (3rd, 4th, 5th, and 6th). She randomly selects 25 students from each grade. Which sampling method is she using?
A Simple random sampling
B Cluster sampling
C Stratified random sampling
D Systematic random sampling

Stratified random sampling divides the population into non-overlapping groups (strata) based on a shared characteristic, then draws a random sample from each stratum. Here, grade level defines the strata and the researcher samples from all four. This differs from cluster sampling (choice B), where only some groups are selected and all members in those groups are included. An SRS (choice A) would not guarantee representation from all grade levels.

Q174. A survey question reads: 'Most experts agree that social media is harmful to teenagers. Do you support legislation restricting social media use for minors?' What is the primary source of bias in this question?
A Nonresponse bias, because some people will refuse to answer
B Undercoverage, because teenagers are not included in the survey
C Response bias from leading question wording
D Voluntary response bias, because the survey is optional

The phrase 'most experts agree' is a leading cue that pressures respondents to agree, producing response bias — inaccurate answers caused by how the question is worded or asked. This inflates the apparent support for restrictions regardless of respondents' true views. Nonresponse bias (choice A) and voluntary response bias (choice D) relate to who participates, not how questions are worded. Undercoverage (choice B) relates to who has a chance to be included in the sample.

Q175. A researcher suspects that age is related to how people respond to a new physical therapy technique. To control for the effect of age, she groups participants into three age brackets (18–35, 36–55, 56+) and randomly assigns participants within each bracket to treatment or control. This experimental design technique is called:
A Stratified random sampling
B Blocking
C Cluster randomization
D Matched-pairs design

Blocking is used in experiments to control for known sources of variability by grouping similar experimental units together (into blocks) before randomizing treatments within each block. This reduces the chance that the variable (age, here) confounds the results. Stratified sampling (choice A) is a sampling technique, not an experimental design technique. Matched-pairs (choice D) specifically pairs two similar subjects, whereas blocking is a broader approach applied here to three groups of more than two.

Q176. A university emails a survey to all 10,000 enrolled students about dining hall quality, and 800 students respond. The administration reports that 72% of students are satisfied. Why is this statistic potentially misleading?
A Email surveys always produce convenience samples with no random mechanism
B Students who have strong opinions — especially those who are dissatisfied — are more likely to respond, introducing nonresponse bias
C The sample size of 800 is too small to represent 10,000 students
D The survey should have used a stratified design by major to be valid

Even though all students were emailed, nonresponse bias is introduced because the 800 who responded likely differ from the 9,200 who did not. Students with strong feelings (positive or negative) are more likely to take time to respond. The 72% figure may not represent the opinions of all 10,000 students. A sample of 800 from 10,000 (choice C) can actually be quite adequate if representative — sample size alone does not determine bias.

Q177. An experiment tests the effect of caffeine on reaction time. Subjects are randomly assigned to receive either a caffeinated drink or an identical-tasting non-caffeinated drink. Subjects do not know which drink they received, but the experimenter who measures reaction times does know. This design is best described as:
A A double-blind experiment
B A single-blind experiment
C An observational cohort study
D A matched-pairs experiment

In a single-blind experiment, one party — typically the subject — does not know whether they are receiving the treatment or control. Here, subjects do not know which drink they received, but the experimenter does, making this single-blind. A double-blind design (choice A) would require that neither the subjects nor the people measuring responses know the assignments, which reduces experimenter bias. Since subjects are randomly assigned to conditions, this is an experiment, not an observational study (choice C).

Q178. A quality control team tests three assembly methods by randomly assigning 10 workers to each method and recording the number of defects per hour. The principle of experimental design that requires applying each method to multiple workers (rather than one worker per method) is:
A Blinding
B Blocking
C Replication
D Control

Replication means applying each treatment to multiple experimental units so that results can be attributed to the treatment rather than chance variation among individuals. With only one worker per method, a single unusually skilled or unskilled worker could dominate the results. Blocking (choice B) would involve grouping workers by a known characteristic (like experience level) before randomizing. Control (choice D) refers to having a baseline group for comparison.

Q179. A researcher is studying the relationship between hours of sleep and GPA. She collects data from 300 college students by asking them to self-report their average nightly sleep and providing their official GPA. She finds that students who sleep more have higher GPAs. What is the most important limitation of this study?
A Self-reported sleep data may be inaccurate, but this affects precision, not the direction of the association
B Because this is an observational study, confounding variables (such as stress levels, course load, or part-time employment) prevent causal conclusions
C The sample size of 300 is insufficient to detect a true relationship between sleep and GPA
D GPA is a categorical variable and should not be used as a response variable in this context

Observational studies cannot establish causation because confounding variables may be responsible for the observed association. Students who sleep more may also have lighter course loads, less financial stress, or better mental health — all of which could raise GPA independently of sleep. Self-reporting inaccuracy (choice A) is a valid concern but is secondary to the fundamental limitation of observational design. GPA is a quantitative variable (choice D), making that choice factually incorrect.

Q180. A researcher surveys adults every 5 years for 20 years to track how their diet changes as they age. This type of study design is best described as a:
A Randomized controlled experiment
B Cross-sectional observational study
C Longitudinal (prospective) observational study
D Case-control study

A longitudinal study follows the same subjects over time, collecting data at multiple points. This is also called a prospective study when subjects are followed forward in time. A cross-sectional study (choice B) collects data from different individuals at a single point in time — for example, comparing the diets of people of different ages simultaneously rather than tracking the same people over decades. Since no treatment is assigned, this is observational, not an experiment (choice A).

Q181. A city planner wants to estimate the average commute time for residents in a large metropolitan area. The planner divides the metro area into 50 zip codes, randomly selects 10 zip codes, and then randomly samples 30 residents from each selected zip code. Which sampling method does this describe?
A Pure cluster sampling, because zip codes are randomly selected
B Stratified random sampling, because residents are sampled within zip codes
C Multistage sampling, combining cluster selection with simple random sampling within clusters
D Systematic sampling, because a fixed number is drawn from each zip code

Multistage sampling involves selecting sampling units in stages, each using some form of random selection. Here, stage 1 randomly selects zip codes (clusters) and stage 2 randomly samples individuals within those zip codes. Pure cluster sampling (choice A) would survey all residents of the selected zip codes, not a random subset. Stratified sampling (choice B) would require sampling from all 50 zip codes, not just 10. Systematic sampling (choice D) uses fixed intervals on a list.

Q182. A television network uses an online poll asking viewers: 'Was tonight's season finale satisfying?' and receives 50,000 responses, with 85% saying yes. A statistician warns that this result is unreliable. Which critique is most statistically valid?
A The sample size of 50,000 is too large and will always produce statistically significant results
B Online polls are always biased because the internet overrepresents younger demographics
C The voluntary response mechanism means only motivated viewers respond, and the large sample size does not compensate for this bias
D The question wording is neutral, so the only concern is undercoverage of non-viewers

A large sample size reduces sampling variability but does not reduce bias from a flawed sampling method. Voluntary response samples systematically overrepresent people with strong opinions, often enthusiasm. The 85% figure reflects whoever chose to respond, not a representative sample of all viewers. This illustrates a key principle: precision (from a large \(n\)) and accuracy (freedom from bias) are independent qualities. A biased poll of 50,000 is worse than an unbiased poll of 1,000.

Q183. In an experiment studying the effect of background music on studying, a researcher uses gender as a blocking variable, randomly assigning males and females separately to music or no-music conditions. What is the primary purpose of incorporating gender as a block?
A To increase the total sample size of the experiment
B To prevent double-counting of subjects who identify as multiple genders
C To reduce variability in the response variable by ensuring gender is equally distributed across treatment groups, allowing a cleaner comparison
D To make the study double-blind with respect to gender

Blocking controls for a known source of variability (gender) by ensuring it is balanced across treatment conditions. If gender affects studying performance independently of music, blocking prevents gender from being a confounding variable. Without blocking, random assignment might by chance place most males in one group, distorting the comparison. Blocking does not increase sample size (choice A) and has nothing to do with blinding (choice D).

Q184. A study finds that students who take notes by hand score higher on conceptual test questions than students who type notes on laptops. The researchers conclude that handwriting causes better conceptual understanding. A critic argues this conclusion is invalid. Which of the following is the strongest methodological basis for the critic's argument?
A The study likely had too small a sample size to detect a true effect
B Students who prefer handwriting may differ systematically from laptop users in ways that also affect conceptual learning, and since no random assignment was used, causation cannot be established
C Conceptual test questions are not a valid measure of learning outcomes in educational research
D The study should have used a matched-pairs design matching students on typing speed

Without random assignment of students to note-taking method, self-selection creates potential confounding. Students who choose to write by hand may differ from laptop users in study habits, learning style, prior academic achievement, or course type — any of which could explain the difference in scores independently of note-taking method. This is the fundamental limitation of observational studies. Random assignment (as in a true experiment) is the only design feature that controls for all confounding variables, including those not yet measured.

Q185. A stratified random sample is drawn from a population with 5 strata of very different sizes. Researchers use proportional allocation, sampling a fraction \(\frac{n_h}{N_h}\) of each stratum. Compared to a simple random sample of the same total size, which statement most accurately describes the advantage of this stratified design?
A The stratified sample eliminates all bias while the SRS does not
B The stratified sample guarantees that the sample mean equals the population mean
C The stratified sample tends to have lower variability in estimates of population parameters when the strata are internally homogeneous but differ from each other
D The stratified sample allows researchers to make causal claims that an SRS does not

Stratified sampling improves precision (reduces variance) when the strata are internally similar but differ from one another — a property measured by low within-stratum variance and high between-stratum variance. By ensuring representation from every stratum, the estimate of the population parameter is more stable across repeated samples compared to an SRS of equal size. However, stratified sampling does not eliminate bias (choice A) — both SRS and stratified sampling are unbiased estimators under proper random selection. Neither design enables causal conclusions (choice D), which requires an experiment.

Q186. A school district tests a new literacy curriculum by randomly selecting 6 schools and implementing the new curriculum in all of them, then comparing outcomes to 6 other schools using the traditional curriculum. A statistician warns that the standard error formula used, which assumes independent observations, will underestimate the true standard error. Why?
A Schools are selected randomly, so independence is guaranteed by design
B Students within the same school share teachers, resources, and environment, making their outcomes correlated — this within-cluster correlation inflates the effective variance beyond what formulas assuming independence would predict
C Comparing only 12 schools violates the large-sample conditions for inference
D The study should have used a within-subjects design to ensure independence

When cluster sampling is used (here, selecting whole schools), individuals within a cluster share common influences, producing intra-cluster correlation. Standard formulas for standard error assume independent observations; when observations within a cluster are positively correlated, the effective sample size is smaller than the actual count of students, meaning the true standard error is larger than the formula suggests. This design effect must be accounted for with multilevel or cluster-robust methods. Random selection of schools (choice A) addresses external validity but does not create independence within schools.

Q187. In a randomized experiment on a new antidepressant, 40% of the placebo group reports significant mood improvement after 8 weeks. A researcher argues this demonstrates that the drug trial is flawed and should be restarted. An experienced statistician disagrees. Which response best reflects correct statistical reasoning?
A The statistician is wrong — a 40% placebo response rate proves the randomization failed
B The statistician is correct — a substantial placebo response is normal and is why blinding and a placebo control are included in the design; the drug's effect is evaluated by comparing treatment and placebo groups, not by evaluating the placebo group alone
C The statistician is correct — 40% response in the placebo group means the study needs a larger sample to detect a drug effect
D The statistician is correct — placebo responders should be removed from the analysis before comparing groups

The placebo effect — whereby patients improve simply from receiving treatment and believing it may help — is a well-documented psychological phenomenon. Double-blind, placebo-controlled trials are specifically designed to isolate the drug's effect above and beyond the placebo effect. A 40% improvement in the placebo group is not a design flaw; it is expected and accounted for. The drug is judged effective only if the treatment group improves significantly more than the placebo group. Removing placebo responders (choice D) would introduce bias by altering group composition after randomization.

Q188. A researcher conducts a randomized controlled experiment and finds a statistically significant result (\(p < 0.01\)). A colleague claims this means the result has high external validity. Which of the following best evaluates this claim?
A The claim is correct because statistical significance directly measures how generalizable results are
B The claim is incorrect — statistical significance (a low \(p\)-value) addresses internal validity (whether the treatment caused the effect in this study), but external validity depends on how the sample was selected and whether it represents the broader population of interest
C The claim is correct because random assignment guarantees both internal and external validity
D The claim is incorrect because \(p < 0.01\) is not significant enough to support any causal or generalizable claims

Internal validity — whether the experimental treatment actually caused the observed effect — is strengthened by random assignment. External validity — whether results generalize to other people, settings, or times — depends on how subjects were recruited and selected, not on the \(p\)-value. A study with a highly significant result from a convenience sample of college undergraduates may have strong internal validity but weak external validity for the general adult population. A small \(p\)-value only tells us the result is unlikely under the null hypothesis; it says nothing about generalizability.

Q189. A consumer research firm wants to estimate the proportion of U.S. adults who would switch to a plant-based diet if prices were equivalent to meat. They use random-digit dialing to phone households between 9 AM and 5 PM on weekdays and get responses from 1,200 adults. Which of the following best identifies the primary bias and its likely direction?
A Voluntary response bias — people who answer the phone self-select, likely underestimating interest in plant-based diets
B Undercoverage bias — adults who work outside the home during weekday business hours are systematically excluded, and these working adults may differ in dietary preferences or price sensitivity from those reached
C Response bias — phone respondents tend to give socially desirable answers, overestimating willingness to adopt healthier diets
D Nonresponse bias — some households called will not answer, but since the sample is random, this does not affect the estimate

Calling households on weekdays between 9 AM and 5 PM systematically excludes people who work outside the home. Stay-at-home caregivers, retirees, and those who work from home are overrepresented. This is undercoverage — a segment of the target population (working adults) had no chance of being included. The direction of bias is uncertain but plausible: income, lifestyle, and food-purchasing behavior likely differ between those home during business hours and those at work. Choice D incorrectly claims random selection eliminates nonresponse bias — it does not.

Q190. An experiment tests two factors: lighting condition (bright vs. dim) and task type (creative vs. analytical). Subjects are randomly assigned to all four combinations. Results show that bright lighting improves analytical task performance but slightly hinders creative task performance, while dim lighting shows the opposite pattern. This result is best described as:
A A main effect of lighting condition on performance
B A confounding relationship between lighting and task type
C An interaction effect between lighting condition and task type
D A blocking effect in which task type controls for lighting variability

An interaction effect occurs when the effect of one factor on the response variable depends on the level of a second factor. Here, lighting does not have a consistent directional effect on performance across all tasks — its effect reverses depending on whether the task is creative or analytical. A main effect (choice A) would mean lighting consistently improves or hinders performance regardless of task type. Confounding (choice B) would mean the factors are not independently controlled, but since subjects are randomly assigned to all four combinations, this is a properly designed factorial experiment.

Q191. A researcher assigns numbers 1 through 500 to all students in a school and uses a random number generator to select 50 students for a survey. Which sampling method is being used?
A Stratified random sample
B Simple random sample
C Cluster sample
D Systematic sample

A simple random sample gives every individual — and every group of individuals of the same size — an equal chance of being selected. Using a random number generator on a numbered list is the textbook implementation of this method. A stratified sample would require dividing students into subgroups (strata) first and sampling within each. A cluster sample would randomly select entire pre-formed groups. A systematic sample would pick every kth individual from the list.

Q192. A radio station asks listeners to call in and vote on whether a new city ordinance should be passed. Of the 800 callers, 91% oppose the ordinance. Which type of bias most threatens the validity of this result?
A Undercoverage bias
B Nonresponse bias
C Voluntary response bias
D Measurement bias

Voluntary response bias occurs when individuals self-select into a sample, typically because they feel strongly about the topic. People who oppose the ordinance are far more motivated to call in than those who mildly support it, skewing results dramatically. Nonresponse bias applies when selected participants fail to respond — here, no one was randomly selected in the first place. Undercoverage bias occurs when part of the population has no chance of being sampled, which is a secondary concern but not the primary flaw in this design.

Q193. In a well-designed experiment testing a new medication, what is the primary purpose of including a control group?
A To eliminate the placebo effect from influencing any subject's response
B To provide a baseline against which the treatment group's results can be compared
C To ensure all subjects receive the same dosage of the experimental medication
D To randomly assign subjects to treatment and non-treatment conditions

The control group provides a baseline — a reference point showing what happens under standard or no-treatment conditions. By comparing the treatment group to the control group, researchers can attribute differences to the treatment rather than to chance or outside factors. The control group does not eliminate the placebo effect; it actually helps measure it, since control subjects often receive a placebo. Random assignment is a design feature of the entire experiment, not a function of the control group specifically.

Q194. A school district wants to survey student attitudes about extracurricular activities. Administrators divide all students into four groups by grade level (9th, 10th, 11th, 12th) and then randomly select 30 students from each grade. Which sampling method does this describe?
A Cluster sampling, because students are organized into pre-existing groups
B Stratified random sampling, because random samples are drawn from each subgroup
C Systematic sampling, because equal numbers are selected from each group
D Simple random sampling, because individual students are selected randomly

Stratified random sampling divides the population into non-overlapping subgroups called strata — here, the four grade levels — and then takes a random sample within each stratum. This ensures representation from every grade. Cluster sampling would involve randomly selecting entire grades and surveying everyone within the chosen grades. The fact that equal numbers (30) are drawn from each stratum is a choice by the researcher, not what defines stratified sampling.

Q195. A researcher notices that neighborhoods with more coffee shops tend to have lower rates of cardiovascular disease and concludes that coffee shop access reduces heart disease risk. What is the most significant flaw in this conclusion?
A The study uses a sample that is too small to establish any relationship
B Observational studies cannot use correlation as a measure of association
C Confounding variables, such as neighborhood income or walkability, may explain both variables
D The researcher should have used a systematic sample rather than observational data

This is a classic confounding variable problem. Wealthier or more walkable neighborhoods may simultaneously attract more coffee shops and have residents with healthier lifestyles, better healthcare access, and lower stress — all of which reduce cardiovascular disease. The presence of coffee shops is not necessarily causing better health outcomes; both may be caused by a third variable. Observational studies can legitimately measure correlation; the flaw is inferring causation from it. Sampling method is not relevant here — the issue is causal reasoning from an observational dataset.

Q196. A health researcher selects every 15th person who checks into a hospital emergency room over a two-week period to complete a patient experience survey. Which sampling method is this, and what is a key concern with it?
A Cluster sampling; entire shifts may be over-represented if clusters are not randomly chosen
B Stratified sampling; patients may not be proportionally distributed across strata
C Systematic sampling; a periodic pattern in arrivals could introduce bias
D Simple random sampling; the sample frame may not include all possible patients

Selecting every kth individual from a list or sequence is systematic sampling. The key concern is periodicity: if hospital arrivals follow a repeating pattern — for example, certain types of patients consistently arrive at the same times of day or days of the week — then every 15th patient may systematically over- or under-represent certain groups, violating the goal of representativeness. This is a known weakness of systematic sampling compared to simple random sampling. Cluster and stratified sampling involve different structures that do not apply here.

Q197. Researchers test a new anti-anxiety medication by randomly assigning 200 participants to receive either the medication or an identical-looking sugar pill. Neither the participants nor the clinicians assessing their anxiety levels know which treatment each person received until the study is complete. Which term best describes this experimental design?
A Single-blind randomized controlled experiment
B Double-blind randomized controlled experiment
C Matched-pairs observational study with placebo control
D Stratified experiment with blinded assessment

A double-blind experiment means that neither the subjects nor the researchers interacting with or evaluating them know who received which treatment. This controls for two sources of bias: the placebo effect (subjects who know they received treatment may report improvement regardless) and researcher bias (clinicians who know the treatment assignment may unconsciously rate outcomes differently). Because subjects are randomly assigned to treatment conditions, this is an experiment, not an observational study. A single-blind design would only keep one party — usually the subject — unaware of the treatment.

Q198. A researcher studies whether a mindfulness app improves academic performance. She suspects that baseline stress levels vary substantially among students and will affect outcomes. She divides students into three blocks based on a pre-study stress assessment (low, moderate, high) and randomly assigns students within each block to the app or a control condition. What is the primary statistical advantage of this block design over a completely randomized design?
A It eliminates confounding due to stress level entirely, allowing a causal conclusion
B It allows the researcher to test whether the app's effectiveness depends on stress level
C It reduces experimental error by ensuring treatment groups are balanced on stress level
D It makes the experiment double-blind with respect to the stress blocking variable

The primary purpose of blocking is to reduce variability — or experimental error — by grouping subjects who are similar with respect to a variable that affects the response (here, stress level). Within each block, subjects are comparable, so differences between treatment and control within a block are more likely due to the app itself. This makes the experiment more sensitive to detecting a real treatment effect. Blocking does not eliminate confounding entirely; it controls for one known source. Testing whether stress level modifies the treatment effect would require a factorial design or interaction analysis, which is a separate question. Blocking has nothing to do with blinding.

Q199. A national polling firm calls randomly selected landline phone numbers on weekday afternoons to estimate the proportion of adults who support a proposed education policy. The poll contacts 1,200 people, of whom 800 complete the survey, and reports a margin of error of \(\pm 3\%\). Which of the following identifies the most serious threat to the validity of this estimate?
A A margin of error of \(\pm 3\%\) is too imprecise to draw conclusions about a national population
B The 400 people who did not complete the survey introduce voluntary response bias
C The sampling frame excludes cell-phone-only households and people away from home on weekdays, creating undercoverage bias
D Calling landline numbers is a form of cluster sampling and cannot produce a valid margin of error

Undercoverage bias arises when part of the population has no chance — or a systematically lower chance — of being included. Adults who rely only on cell phones (disproportionately younger people), those who work outside the home on weekday afternoons, and those without phone access are all excluded from this sampling frame. If these groups hold systematically different opinions about education policy, the estimate will be biased in ways that the margin of error — which only reflects sampling variability, not bias — cannot correct. The 400 non-completers represent nonresponse bias, not voluntary response bias. A \(\pm 3\%\) margin of error is a standard and acceptable level for national polls.

Q200. A study randomly assigns 120 employees to one of two workplace wellness programs (Program X and Program Y) for 12 weeks and measures productivity gains. Program X shows a statistically significant advantage. A methodologist raises the concern that both groups may have improved mainly due to the Hawthorne effect rather than the content of either program. If the Hawthorne effect is the dominant factor, what would be the expected pattern of results, and what design feature would best address this concern?
A Both groups would show little improvement; adding a third group receiving a more intensive intervention would address it
B Program X would show large gains and Program Y would show none; a matched-pairs design would address it
C Both groups would show substantial improvement relative to a true control, and the difference between programs would be small; a no-treatment control group would help isolate the Hawthorne effect
D Program Y would outperform Program X because less-studied groups respond more naturally; stratifying by department would address it

The Hawthorne effect refers to the tendency for people to change their behavior — usually improving it — simply because they know they are being observed or are part of a study, regardless of the specific treatment. If this effect dominates, both groups would show substantial productivity gains relative to their baseline, but the difference between Program X and Program Y would be modest or negligible, since both groups are being observed equally. The best design fix is including a true no-treatment control group that is also observed — this allows researchers to estimate how much improvement is attributable to being studied versus to the actual wellness content. A matched-pairs design addresses between-subject variability, not observation effects.

Study tip

Focus on understanding.

Focus on understanding core concepts before memorizing details. Use the game modes to test yourself repeatedly — spaced repetition is proven to boost long-term retention.

Up next

Related units

Quick summary

This unit covers sampling methods, observational studies, experiments and bias — essential concepts for AP Statistics. Use our interactive study games to test your understanding, or review questions in traditional format below.

Key concepts
  • Sampling methods
  • Observational studies
  • Experiments
  • Bias
What you need to know

Key Concepts Breakdown

1 Sampling Methods

Students must be able to identify and distinguish between simple random sampling (SRS), stratified random sampling, cluster sampling, systematic sampling, and convenience sampling. The exam tests whether students can recognize which method was used in a scenario and explain why a method does or does not produce a representative sample. Understanding that only probability-based sampling methods allow valid inference to the population is critical.

Key Points

  • SRS: every individual and every group of size n has an equal chance of selection — use a random number table or generator
  • Stratified: population divided into homogeneous groups (strata), then SRS taken from each; reduces variability
  • Cluster: population divided into heterogeneous groups (clusters), then entire clusters randomly selected; used for practicality
  • Convenience and voluntary response samples are biased and cannot support inference — always identify these as flawed
Example

A principal wants to survey 60 students about cafeteria food. She divides the school into grade levels (9, 10, 11, 12) and randomly selects 15 students from each grade. What sampling method is this, and what is its advantage over SRS?

Explanation

This is stratified random sampling because the population is divided into subgroups (grades) and an SRS is drawn from each stratum. The advantage is that it guarantees representation from every grade level, which reduces sampling variability compared to SRS, which might by chance under-represent a grade. On the exam, you must name the method AND justify why it is appropriate or advantageous.

2 Observational Studies

In an observational study, researchers measure variables without imposing any treatment — they simply observe. Because no random assignment occurs, observational studies cannot establish causation; they can only identify associations. The exam frequently asks students to distinguish observational studies from experiments and to explain why causation cannot be concluded.

Key Points

  • No manipulation of variables — researcher observes and records naturally occurring behavior
  • Confounding variables (lurking variables) are always a threat; an observed association may be explained by a third variable
  • Retrospective studies look backward at existing records; prospective studies follow subjects forward in time — both are observational
  • Correct language: 'there is an association between X and Y,' never 'X causes Y' for observational data
Example

Researchers examine hospital records and find that patients who received more frequent nurse check-ins had shorter hospital stays. A student concludes that frequent check-ins cause faster recovery. Is this conclusion justified?

Explanation

No — this is an observational study because no treatment was randomly assigned; the researchers simply examined existing records. The conclusion of causation is not justified because confounding variables exist: for example, less severely ill patients may naturally recover faster AND receive more routine check-ins. On the exam, you must identify the study type, name a plausible confounding variable, and use the word 'association' rather than 'causation.'

3 Experiments

A well-designed experiment must include random assignment of treatments to experimental units, which is the only design feature that allows a cause-and-effect conclusion. Students must know the three principles of experimental design — control, randomization, and replication — and must be able to identify and explain control groups, placebo effects, blinding, and blocking. The exam will ask students to design an experiment or critique a flawed one.

Key Points

  • Random assignment (not random sampling) is what allows causal inference — it balances confounding variables across treatment groups
  • Control group provides a baseline; placebo controls for the psychological effect of receiving treatment
  • Double-blind: neither subjects nor evaluators know treatment assignment — eliminates response bias and evaluator bias
  • Blocking: group experimental units by a known source of variability (e.g., sex, age) before random assignment — reduces variability within blocks, analogous to stratification in sampling
Example

A researcher wants to test whether a new study app improves SAT math scores. She randomly assigns 50 students to use the app for 8 weeks and 50 students to use no supplemental tool. Both groups take a pre-test and post-test. Identify the experimental units, explanatory variable, response variable, and one improvement to the design.

Explanation

Experimental units are the 50+50 students; the explanatory variable is app use (app vs. no app); the response variable is change in SAT math score. One improvement would be to use a placebo — have the control group use a 'sham' app with no educational content — so that any improvement from merely using an app (Hawthorne effect) is controlled. Alternatively, blocking by initial math ability would reduce variability. Exams often award points for identifying the improvement AND explaining why it strengthens the study.

4 Bias

Bias is a systematic tendency for a sample statistic to over- or underestimate the population parameter. Students must be able to identify sources of bias by name, explain the direction of the bias (will responses be too high or too low?), and distinguish bias from variability. Increasing sample size does not reduce bias — only changing the design does.

Key Points

  • Sampling bias: non-probability samples (voluntary response, convenience) systematically exclude parts of the population
  • Response bias: question wording, social desirability, or interviewer presence causes subjects to answer inaccurately
  • Undercoverage bias: some groups in the population have little or no chance of being selected (e.g., online survey excludes those without internet)
  • Nonresponse bias: individuals selected for the sample who do not respond likely differ systematically from those who do — cannot be fixed by selecting more people
Example

A magazine publishes an online poll asking readers: 'Do you agree that violent video games are destroying America's youth?' Of the 10,000 people who responded, 84% agreed. Identify two sources of bias and explain how each affects the results.

Explanation

First, voluntary response bias: only readers who feel strongly (likely those who agree) bother to respond, inflating the percentage who agree. Second, response bias from question wording: the loaded phrase 'destroying America's youth' nudges respondents toward agreement, further inflating the 'agree' percentage. Both biases push the estimate in the same direction — upward — making 84% a severe overestimate of the true population proportion. On the exam, always state the name of the bias, identify the mechanism, and specify the direction of distortion.

FAQ

Questions, answered.

What is Collecting Data?

Collecting Data is Unit 3 of AP Statistics, covering sampling methods, observational studies, experiments and bias.

How to study for AP Statistics Unit 3?

Start with the Quick Summary above, review the Key Concepts, then test yourself with our interactive study games. Aim for 80%+ accuracy before moving on.

How many questions are in this unit?

This unit has 200 review questions, each with a written explanation, playable across 5 different game modes or readable in plain-text mode.