Ask most people which is better, data you collected yourself or data taken from someone else, and the instinct answers before the economics does: surely your own. The instinct is wrong often enough to be dangerous. The distinction between primary vs secondary data is about a researcher’s relationship to the collection, not about quality: primary data are gathered by the researcher for the question at hand, through surveys, experiments, and interviews they designed; secondary data were collected by someone else, for some other purpose, and are being reused, which describes the national accounts, the household surveys, the administrative records, and nearly every series in every statistical database economists live in. Empirical economics is unusual among sciences in running overwhelmingly on secondary data, and the craft that makes this respectable is not collection technique but source criticism: knowing, for data one did not gather, exactly how they came to exist. The distinction matters because each kind fails differently, and because the choice between them is a real decision with prices on both sides.
What Each Kind Buys, and What Each Costs
Primary data buy fit. The researcher controls the definitions, the population, the timing, and the questions, so the data can match the research question exactly: the precise concept of informality the study needs, the specific firms, the outcome nobody else measures. The price is everything else: collection is slow and expensive, samples are small because budgets are finite, and the collector inherits every methodological risk personally, from questionnaire wording to refusals, the full obstacle course described in our guides to data collection in economics and survey design. The purest primary form in modern economics is the experiment run in the world itself, the practice covered in our article on field experiments, where the data are generated, not merely gathered, under the researcher’s design.
Secondary data buy scale, history, and comparability. No dissertation budget could reproduce a national labor force survey, a century of price indices, or trade records covering every country; reusing them grants a lone researcher the field apparatus of statistical agencies at zero collection cost, with samples and time spans primary work cannot approach. The price is the mirror image of fit: the definitions are someone else’s, chosen for someone else’s purpose. The unemployment concept is the agency’s, not the study’s; the income variable was built for tax administration, not welfare measurement; the category boundaries fall where a ministry needed them decades ago. Secondary work is therefore a permanent exercise in translation, checking how far the inherited measure is from the concept the question actually needs, and honest papers document the gap instead of assuming it away.
The Craft of Reuse: Source Criticism
Because economics runs on reused data, its distinctive skill is interrogating a series one did not build, and the checklist in the figure is the working version. Who collected it, and with what incentives: a statistical agency under political pressure, a ministry reporting its own performance, and an industry association counting its members are three different witnesses. For what purpose: administrative data record what the administration needed, which is why tax records see income the tax code sees and miss what it exempts. How are the concepts defined, and do the definitions match the question: “unemployed”, “firm”, and “income” each carry official definitions that diverge from their everyday and theoretical meanings. What changed over time: methodologies are revised, base years shift, and category boundaries move, so a fifty-year series is rarely one series, which is one reason the structural shifts examined in our article on structural breaks are sometimes artifacts of measurement rather than events in the economy. And how is it revised: early releases of GDP and employment are estimates that get corrected for years, so a study of decisions made in real time must use the data as decision-makers saw them, not the polished final vintage. None of these checks requires touching a questionnaire; together they are the difference between reusing evidence and repeating someone’s numbers.
Choosing, and the Compilation Trap
The choice in practice is sequential rather than ideological. The literature-and-data reconnaissance that begins any project, the discipline of a systematic literature review, establishes what secondary sources already cover; primary collection is reserved for the residual, the concept nobody measures, the population no frame reaches, the mechanism no record captures, and strong applied work frequently layers a small primary instrument over a large secondary base, the combination logic treated in our article on mixed method research. One boundary deserves special care in the database age: the distinction between a source and a compilation. The great international databases are aggregators, invaluable for access and comparability, but they stand one step further from the collection than the national agencies whose numbers they gather, and details, definitions, breaks, footnotes, can be lost in transmission. The working rule is to cite and check the original producer wherever it matters, treat compilations as catalogs rather than witnesses, and treat data whose provenance cannot be established at all, re-hosted files of unknown ancestry, as leads to a source rather than sources themselves. Provenance is to data what identification is to estimates: the thing the number cannot supply about itself.
MASEconomics Explains
3 economic concepts behind primary vs secondary data
These concepts are explored in depth across our educational articles library.
Explore the MASEconomics BlogConclusion
The primary vs secondary data distinction classifies relationships, not quality: data gathered by the researcher for the question, against data inherited from collectors who had other purposes. Each buys what the other cannot. Primary collection purchases exact conceptual fit at the price of cost, scale, and personally inherited risk; secondary reuse purchases the scale, history, and comparability of the world’s statistical apparatus at the price of living inside someone else’s definitions. Economics, run largely on the second kind, compensates with a discipline the first kind never needed: source criticism, the systematic interrogation of who made the numbers, why, under what definitions, and with what revisions.
The practical order follows from the prices. Exhaust the secondary landscape first, since duplication of what agencies already measure wastes any budget; collect primarily for the genuine residual; and wherever reuse occurs, trace provenance to the original producer, treating compilations as catalogs and unprovenanced files as rumours. The instinct that one’s own data are better dissolves into something more useful: the recognition that all data are somebody’s collection decisions frozen into numbers, and that knowing whose decisions, and why, is the part of the evidence no download includes.
Frequently Asked Questions
What is the difference between primary and secondary data?
Primary data are collected by the researcher specifically for the current question, through surveys, experiments, or interviews they designed. Secondary data were collected by someone else, usually for a different purpose, and are being reused: official statistics, administrative records, and existing survey archives. The distinction concerns the relationship to collection, not quality.
Is primary data better than secondary data?
Neither is better; they buy different things. Primary data offer exact fit to the research concept at high cost and small scale; secondary data offer scale, long history, and cross-country comparability at the price of inherited definitions. Most empirical economics runs on secondary data, disciplined by careful source criticism, with primary collection reserved for what nobody else measures.
Are databases like national statistics portals primary or secondary sources?
For the reusing researcher they are secondary data, and many are additionally compilations: aggregators standing one step further from collection than the national agencies whose numbers they gather. Good practice cites and checks the original producer for anything that matters, since definitions and breaks can be lost in transmission.
When should a researcher collect primary data?
When the question requires a concept, population, or mechanism that no existing source measures: a definition of informality no survey uses, firms no registry lists, beliefs no questionnaire asks about. The decision comes after reconnaissance of the secondary landscape, because duplicating what agencies already measure spends scarce budget on worse versions of existing data.
What questions should be asked of any secondary dataset?
Five reliably: who collected it and with what incentives; for what purpose the records exist; how the key concepts are defined and whether those definitions match the question; what methodological changes interrupt the series over time; and how the data are revised, so that real-time studies use the vintage contemporaries actually saw.
Thanks for reading! All data are somebody’s decisions frozen into numbers; the download never includes whose. Happy learning with MASEconomics