The unemployment rate that moves markets every month is not a count of the unemployed. It is an inference from interviews with a few tens of thousands of households, standing in for a country of hundreds of millions, and the reason such a small slice can speak for the whole is not luck but design. Sampling methods are the procedures that decide which slice of a population gets observed, and they carry more weight than almost any other methodological choice, because no analysis, however sophisticated, can repair a sample that was selected badly: the flaw enters before the first number is recorded. The three workhorse designs, simple random, stratified, and cluster sampling, exist because the textbook ideal collides with three practical problems, and each design solves one of them at a known price. Knowing the designs, and the sins they guard against, is knowing how much to trust every statistic built on them, which is most of the statistics there are.
Random Selection, and Why It Is the Benchmark
Simple random sampling gives every member of the population an equal chance of selection, and that equality is the entire magic. Because chance does the choosing, no characteristic, income, region, opinion, firm size, can systematically over- or under-enter the sample, so the sample resembles the population in expectation on every dimension at once, including the ones nobody thought to list. Randomness also licenses the mathematics: the sampling distributions behind every margin of error and every test in our guide to hypothesis testing assume chance-based selection, and they can even be reconstructed empirically by resampling, the logic of bootstrap methods. The design’s obstacle is mundane: it requires a sampling frame, a complete list of the population to draw from, and for most populations of interest, informal workers, small firms, households in countries with thin registries, no such list exists. The frame problem, not the mathematics, is why the pure design is rare in practice and why the other two exist.
Stratified and Cluster: Two Prices, Two Purchases
Stratified sampling divides the population into layers that matter, regions, firm-size classes, income bands, and draws randomly within every layer. The purchase is twofold. Precision improves, because guaranteeing each stratum its correct share removes one source of sample-to-sample luck; and subgroups are protected, because a small but important group, large firms, a minority region, that simple randomness might catch thinly is guaranteed its quota. Statistical agencies stratify nearly everything for exactly these reasons, oversampling small strata deliberately and reweighting afterwards so the totals still represent the population. The price is informational: stratification requires knowing the layering variable for the whole frame before sampling, which is why it works best where registries are rich.
Cluster sampling solves the opposite problem: no frame and no travel budget. Instead of listing every household in a country, list its villages and city blocks, draw a random set of those clusters, and survey within the chosen ones; frames of villages exist where frames of households do not, and interviewers visit dozens of neighborhoods rather than thousands of scattered addresses. The price is precision. People within a cluster resemble each other, sharing labor markets, schools, and shocks, so each additional respondent from the same village adds less new information than a fresh independent draw would; a clustered sample of ten thousand may carry the informational content of a much smaller random one, a discount measured by the design effect, and the same within-group correlation is why modern regression work computes clustered standard errors when data arrive in groups. Real national surveys, the labor force and household surveys behind the statistics economists consume, are multistage composites: clusters drawn first for feasibility, stratified by region for precision, with random selection inside, the architecture described in our guides to data collection in economics and survey design.
The Sins the Designs Exist to Prevent
Every design above shares one property: chance, not convenience and not choice, decides who enters. The alternative is the convenience sample, whoever is easiest to reach, and its modern industrial form, the online opt-in poll, where respondents select themselves. Self-selection is fatal in a way small samples are not, because the act of volunteering correlates with the thing being measured: the angriest customers review, the most engaged citizens answer, the survivors fill the dataset of firms. No sample size cures it, since collecting more of a biased stream yields a larger biased sample, and the twentieth century’s most famous polling disasters came from enormous convenience samples beaten by small random ones. Related sins are quieter: coverage error, when the frame misses part of the population, phone surveys missing the phoneless, business registries missing the informal economy that looms large in developing countries; and nonresponse, when the randomly selected cannot be found or decline, and those who answer differ from those who do not, the erosion that weighting and follow-up protocols exist to limit, and that has grown into the survey world’s chief modern ailment as response rates fall. The reader’s defense is one question asked of every statistic: who could have entered this sample, and who never could? A number that cannot answer, as our introduction to econometrics frames it, is not yet evidence.
MASEconomics Explains
3 economic concepts behind sampling methods
These concepts are explored in depth across our educational articles library.
Explore the MASEconomics BlogConclusion
Sampling methods decide what a statistic can honestly claim before any analysis touches the data. Simple random selection is the benchmark because chance-based entry makes the sample resemble the population on every dimension at once, and its two great modifications are purchases against practical constraints: stratification buys precision and subgroup guarantees where registries are rich, clustering buys feasibility where frames and budgets are thin, at a precision discount the design effect prices. The national surveys behind the unemployment rate, the price indices, and the household statistics are engineered composites of all three, which is why a few tens of thousands of interviews can speak for hundreds of millions.
The designs also define the sins. Convenience and self-selection break the chance principle and cannot be cured by volume; coverage gaps silently exclude parts of the population from possibility; nonresponse erodes randomness after the draw. For the consumer of statistics the lesson compresses to one habit: before asking what a number says, ask who could have entered the sample behind it, and by what mechanism. Numbers inherit the honesty of their selection, and the selection happened before anyone began to count.
Frequently Asked Questions
Why do statisticians sample instead of counting everyone?
Cost, speed, and, surprisingly, accuracy. A census of millions invites processing errors and stale results by the time it is done, while a well-designed sample of tens of thousands can be interviewed carefully, quickly, and repeatedly, with a margin of error that is known and small. Most official monthly statistics are samples for exactly these reasons.
What is the difference between simple random and stratified sampling?
Simple random sampling draws from the whole population with equal chance, leaving subgroup coverage to luck. Stratified sampling first divides the population into layers, regions, size classes, income bands, and draws randomly within each, guaranteeing every layer its share, improving precision, and allowing small important groups to be oversampled and reweighted.
What is cluster sampling and why accept its precision loss?
It draws whole groups, villages, city blocks, schools, and surveys within the chosen ones. It exists because complete lists of individuals often do not, while lists of places do, and because visiting a few hundred clusters costs far less than reaching scattered addresses. The price is that similar neighbors add less information, a discount the design effect measures.
Why are online opt-in polls unreliable?
Because respondents choose themselves, and the willingness to volunteer correlates with the opinions and traits being measured. The resulting bias does not shrink as the sample grows; a million self-selected responses are a larger biased sample, not a better one, which is why small random samples have historically beaten enormous convenience ones.
Does a sample need to be a large share of the population?
No, and this is the most counterintuitive fact in sampling: precision depends on the absolute size and design of the sample, hardly at all on the fraction of the population covered. A properly random sample of a few thousand describes a country of fifty million about as well as it describes one of five hundred million.
Thanks for reading! Numbers inherit the honesty of their selection, and the selection happens before anyone counts. Happy learning with MASEconomics