Collection of Data
This chapter teaches you how economists and statisticians collect raw data — the fuel for every graph, policy, and business decision — and how to judge whether data you encounter is reliable.
Whether you go into CA, B.Com, civil services, or start a business, you will constantly read reports, surveys, and statistics — this chapter gives you the critical eye to know when to trust a number and when to question it.
Concept
Lots of students think…
"A larger sample is always more reliable than a smaller one — the bigger the better."
Actually…
What matters most is how the sample was chosen, not just its size. A randomly selected 500 households nationwide can represent India better than a biased 5,000 from one city. Selection method beats sample size every time.
By the end of this, you will understand how economists actually collect the numbers you see in news headlines — and how to tell good data from bad data.
What is data and why collect it?
Data is just collected information — numbers, facts, answers to questions. Before any economist, government, or business can make a decision, they need reliable data. Without it, every plan is just a guess.
Before deciding where to open a new government hospital in Kerala, the state health department first collects data: how many people live in each district, how far they travel to reach the nearest hospital, and how many patients each hospital currently handles. Without that data, they might build the new hospital in the wrong place.
Primary data — fresh from the source
Primary data is information you collect yourself, directly, for your specific question. You go out, ask people, measure things, or run experiments. It is fresh and fits exactly what you need — but it takes time and money.
Riya runs a small tiffin service in Pune. She wants to know if office workers in her area would pay ₹80 for a healthy meal. She visits three offices and surveys 60 workers herself — asking them face-to-face. Those responses are primary data: she collected them firsthand for her exact question.
Secondary data — already collected, ready to use
Secondary data is information someone else already collected, usually for a different purpose. You access it from reports, government publications, old surveys, or websites. It is cheaper and faster, but you must check: is it recent? Does it really fit your question?
Riya also downloads the Annual Report on Household Food Expenditure published by the National Statistical Office — it covers lakhs of households across India. She did not collect this; the government did. She uses it to see the national trend in food spending. That is secondary data at work.
Census — counting everyone
A census means you collect data from every single person or unit in the group you are studying. You do not miss anyone. It gives the most complete picture possible, but it is hugely expensive and slow.
India's national Census is conducted every ten years by the government. In 2011, it counted every one of India's 121 crore people — their age, education, occupation, and housing. It cost over ₹2,200 crore and took thousands of government workers years to complete. That is what a true census looks like.
Sampling — studying a small group to understand the big one
Sampling means you pick a smaller group (called a sample) from the larger group (called the population), study the sample carefully, and use those findings to describe the whole population. The key is that the sample must be representative — it must reflect the variety within the full group.
India has about 28 crore households. The NSO cannot visit all of them every year. So it randomly selects around 1 lakh (1,00,000) households — that is just 0.036% of all households — studies them carefully, and uses those findings to estimate nationwide poverty, employment, and food spending. The results shape policies like MGNREGA and the Public Distribution System.
Bias — the silent enemy of good data
Bias means your sample is not truly representative — some groups are left out or over-represented, so your conclusion is skewed. Bias can come from how you choose your sample, how you word your questions, or even how the interviewer behaves.
Suppose a biscuit company surveys customers only inside premium supermarkets in South Mumbai to find out how much Indians spend on biscuits. They will get answers only from wealthy city shoppers — not from rural families, small-town buyers, or low-income households. Their data will show Indians spend far more on biscuits than they actually do. That is sampling bias.
The NSO — India's data backbone
The National Statistical Office (NSO) is the government body that runs India's major surveys. It was earlier known as the NSSO (National Sample Survey Office). It collects data on household income, employment, health, farming, and more — the numbers that shape national policy.
When the government decided to set the minimum wage under MGNREGA, it used NSO survey data on what rural households actually spend on food and basic needs. Without that data, officials would have been guessing. The NSO survey had visited lakhs of rural households across every state to gather those figures.
Notes
The full picture
Every number you see in a newspaper headline — India's GDP grew 7%, unemployment fell to 8%, or rice production hit a record — had to be collected by someone, somewhere. Before any analysis or policy decision happens, someone must gather information in a systematic way. This chapter is about that first and most important step: how data gets collected, and why the method used determines whether we can trust the conclusions.
Data collected firsthand for a specific purpose is called primary data. If a market researcher visits fifty Kirana shops in Patna to record their monthly sales, those records are primary data. Primary data fits your exact question, but it takes time and money. Secondary data is information that already exists — collected earlier, usually for a different purpose. Examples include the Annual Survey of Industries published by the government, or a bank's published balance sheet. Secondary data is cheaper and faster to access, but you must always ask: is this data recent enough? Does it match exactly what I need? A student researching urban poverty today cannot rely on a 2011 census figure and assume nothing has changed.
A census covers every single unit of the population being studied. India's national Census, conducted every ten years, counts every person and household in the country. It is the most complete picture possible, but it is enormously expensive — the 2011 Census cost over ₹2,200 crore and took years. Most data collection, therefore, uses sampling: you select a representative subset of the population, study that subset carefully, and then use its findings to describe the whole. The key word is representative. A sample must reflect the diversity of the full population, or your conclusions will be wrong.
India's National Statistical Office (NSO), which includes what was formerly known as the NSSO (National Sample Survey Office), runs large-scale sample surveys to measure household consumption, employment, poverty, and health. In a typical survey round, around 1 lakh (100,000) households are surveyed out of roughly 28 crore (280 million) households in India — that is about 0.036%. Yet these surveys shape national policy on minimum wages, food subsidies, and social schemes like MGNREGA. The reason such a small fraction can represent the whole country is random selection: every household must have an equal chance of being picked, which prevents any group from being systematically ignored.
The biggest danger in sampling is bias: the sample is not truly representative. If you survey only mobile-phone users to study rural income, you exclude the poorest households who cannot afford phones. If a biscuit company surveys customers only at premium supermarkets in Mumbai, it will overestimate how much Indians spend on snacks. Bias can creep in through how samples are chosen, how questions are worded, or even how interviewers behave. Good data collection design anticipates and controls for all of these. This is why trained enumerators, carefully worded questionnaires, and random selection methods matter so much — not just technical details, but the difference between reliable evidence and misleading statistics.
An Indian example
In 2022, Meera Pillai launched a small organic spice brand and wanted to sell across India. Before spending ₹5 lakh on packaging and logistics, she needed to know: do urban consumers in north India actually buy organic spices online? She could not afford to survey the entire country (a census approach), so she used a sample — 400 households across Delhi, Pune, and Ahmedabad, selected randomly by a market research firm at a cost of ₹40,000. The survey showed 68% were willing to pay ₹120–150 for a 100g organic spice pack, versus ₹60 at a regular shop. She also consulted secondary data — an industry report on organic food sales growth in metro cities — to cross-check the trend. Together, the primary survey (fresh, specific to her product) and the secondary report (broad market context) gave her the confidence to launch. By year two, her brand was stocked in 12 cities. Had she skipped the primary data and relied only on a national report, she would have missed the price sensitivity unique to her specific product category.
Key concepts covered
- Primary vs secondary data
- Sampling vs census
- NSSO & sources
Common misconceptions to watch for
- Wrong belief: Primary data is always better than secondary data because it is newer and more specific. Correction: Both types have their place. Secondary data from a reliable source like the RBI or NSSO can be highly accurate and perfectly adequate for many research questions. The choice depends on what you need, your timeline, and your budget — not a blanket rule.
- Wrong belief: A larger sample is always more reliable than a smaller one. Correction: What matters most is how the sample was chosen, not just its size. A randomly selected sample of 500 households nationwide can represent India better than a biased sample of 5,000 people drawn only from one city or one income group. Selection method beats sample size.
- Wrong belief: The NSO/NSSO surveys all households in India, so its data is like a census. Correction: The NSO conducts sample surveys — typically around 1 lakh households per round, not all 28 crore. Its power comes from random sampling design, not complete coverage. A census and a sample survey are distinct methods with different purposes and costs.
Questions
A restaurant chain needs to decide on a new vegan menu within 2 weeks. Option A: primary survey of 5,000 customers (₹2,50,000, 4 weeks). Option B: secondary national food report (₹25,000, available now, 1 year old). Option C: quick primary survey of 500 customers at top outlets (₹15,000, 3 days). Which is best?
- 1Clarify what information is needed.The restaurant must understand if its own customers want vegan options. This requires data from the relevant customer group, not the general population.
Question 1 of 5 · medium
A government agency studies PM-MUDRA's impact on rural businesses within 6 months. Should it use secondary data (old reports) or conduct a primary survey?
Quiz
Test yourself — pick an answer, then hit "Check" to see the explanation and your running score.
Question 1 of 5 · medium
A government agency studies PM-MUDRA's impact on rural businesses within 6 months. Should it use secondary data (old reports) or conduct a primary survey?
Spotted an arithmetic error or unclear explanation? Suggest an edit — we fix things fast.