Collection of Data
This chapter teaches you where economic data comes from and how to collect it — skills that sit behind every government policy, every business decision, and every exam question that asks you to 'explain with examples'.
Every career that touches numbers — CA, B.Com, economics research, civil services, business — requires you to judge whether data is trustworthy; getting this chapter right also unlocks the rest of Statistics for Economics, because all analysis depends on well-collected data.
Concept
Lots of students think…
"A bigger sample is always more accurate, so a survey of 4,000 people is always better than one of 400."
Actually…
What matters is how you select, not just how many you select. A randomly chosen sample of 400 beats a carelessly chosen sample of 4,000 because random selection removes bias.
By the end of this, you will understand where economic data actually comes from and why collecting it the right way matters more than collecting a lot of it.
What Is Data?
Data is simply collected facts or measurements about something. Before any economist, government, or business can say anything meaningful, someone has to gather the raw information first. Without data, every statement about the economy is just a guess.
When you hear 'Kerala has the highest literacy rate in India,' that claim comes from the Census of India — millions of answers collected, counted, and turned into a number. Without that data collection, we would have no way to know.
Primary vs Secondary Data
Primary data is information you collect yourself, fresh, for your own question. Secondary data is information someone else already collected — you just find and use it. Neither is automatically better; the right choice depends on your question, your deadline, and your budget.
If you go around your class and ask 30 classmates how many hours they study each day, that is primary data — you gathered it yourself. If you use the Kerala State Planning Board's report on student performance in Plus One exams, that is secondary data — someone else already did the work.
Census: Counting Everyone
A census means you survey every single person or unit in the group you are studying — no one is left out. It gives the most complete picture possible, but it is very expensive and takes a long time.
The Census of India, held every ten years, tries to record the age, education, occupation, and housing of all 1.4 billion residents. It costs thousands of crores of rupees and takes years to process — but it gives the government data to plan schools, hospitals, and roads for every corner of the country.
Sampling: Study a Few, Learn About All
Sampling means choosing a smaller, carefully selected group and studying only them. If the group is chosen well, their answers represent the whole population — at a fraction of the cost and time of a census.
Suppose a district in Kerala wants to know the average monthly food spending of households. Instead of visiting every family (lakhs of them), they survey 500 randomly chosen families. Done carefully, those 500 give a reliable answer — and the whole thing costs a few lakhs, not thousands of crores.
How to Collect Primary Data
Once you decide to gather primary data, you need a method. The main options are: written questionnaires people fill in themselves, personal interviews where you ask questions face to face, telephone or online surveys, and direct observation where you watch and record without asking. Each method suits different situations.
Ananya runs a kirana store in Thrissur and wants to know if her customers are buying less packaged atta. She designs a ten-question paper questionnaire and hands it to 40 regular customers over two days. This personal, written method costs almost nothing and gives her exactly the local, current answer she needs — in 48 hours.
Where to Find Secondary Data
Secondary data is everywhere if you know where to look. Government agencies, research institutions, and newspapers all publish data regularly. The key skill is knowing which source to trust and whether it is still fresh enough to be useful.
The Reserve Bank of India (RBI) publishes quarterly reports on money, credit, and inflation. The National Statistical Office (NSO) runs regular surveys on household spending and employment. If you are writing an economics project and you copy a table from the NSO website, you are using secondary data — it already existed, so you just found it.
Bias: The Enemy of Good Data
Bias means your data collection process systematically favours some people or answers over others — so your result is skewed, not true. A bigger sample does not fix bias; only better selection does. Always ask: who collected this data, when, how, and for what purpose?
Say a company wants to measure customer satisfaction and calls only people who bought online. They completely miss customers who shop in their physical stores — those customers might have very different experiences. The result looks positive but is misleading. A leading question like 'Don't you find our service excellent?' has the same problem — it pushes people toward saying yes.
Notes
The full picture
Every economic statement you have ever heard — 'India's unemployment rate rose', 'onion prices doubled', 'Kerala has the highest literacy in India' — is built on data. Data is simply collected facts or measurements about something. Before an economist can analyse anything, someone has to gather the information first. This chapter is about that gathering step: what it means, how it is done, and why doing it badly ruins the entire chain of reasoning that follows.
Data has two broad origins. Primary data is information you collect yourself, fresh, for your specific question. Secondary data is information that was already gathered by someone else — government agencies, research institutions, newspapers — and is now available to you. Think of it this way: if you survey students in your class about how many hours they study, that is primary data. If you use the Kerala State Planning Board's report on student performance, that is secondary data. Both are useful; neither is automatically better. The right choice depends on what you need, when you need it, and how much it costs.
Primary data can be collected in three main ways. The first is a census, also called complete enumeration: you survey every single unit in the group you are studying. The Census of India, conducted every ten years, tries to record details about all 1.4 billion residents — their age, education, occupation, and housing. It is the most complete picture possible, but it costs thousands of crores of rupees and takes years to process. The second method is sampling: you choose a smaller, representative sub-group and study only them. If a district wants to know the average monthly expenditure of households, surveying 500 carefully chosen families can give a reliable answer at a fraction of the cost of visiting every household. The third method covers how you actually ask questions — through written questionnaires, personal interviews, telephone calls, or direct observation. Each approach carries its own strengths and weaknesses that you must weigh before choosing.
Secondary data is everywhere if you know where to look. The Reserve Bank of India (RBI) publishes quarterly reports on credit, money supply, and inflation. The National Statistical Office (NSO), which absorbed the former National Sample Survey Organisation (NSSO), runs regular surveys on household consumption, employment, and income. The Department of Economics and Statistics, Kerala publishes detailed state-level data on agriculture, population, and education. Private rating agencies like CRISIL and ICRA track industries and companies. When you write an economics project and cite a government table, you are using secondary data. The great advantage is speed and cost: the data already exists, so you just download or find it. The danger is that it may be outdated, collected for a different purpose than yours, or contain errors you did not notice.
Quality matters more than quantity in data collection. A small sample chosen carefully through random selection can be more reliable than a large survey contaminated by bias. Bias creeps in when the selection process systematically favours certain people or answers over others. For example, if a company surveys customer satisfaction only by calling people who purchased online, it will miss the experience of customers who buy in shops — skewing the result. Similarly, a leading question like 'Don't you find our service excellent?' pushes respondents toward a positive answer. When you evaluate any data source — whether for your exam, a project, or life — always ask: who collected this, when, how, and for what purpose? Those four questions are the foundation of critical thinking about statistics.
An Indian example
Ananya runs a small kirana store in Thrissur. She hears that sales of packaged atta have fallen across Kerala and wants to know if she should reduce her stock order. She finds two options: the Ministry of Consumer Affairs publishes average wholesale prices, but the latest figures are three months old and cover pan-India trends. Alternatively, she designs a quick questionnaire — ten questions, given to 40 regular customers over two days — asking how many kg of atta they buy per month and whether they switched to a local mill brand recently. Her 40-customer survey (primary data) costs her almost nothing and takes two days. It tells her exactly what her customers in her neighbourhood are doing right now. The government report (secondary data) is free but too old and too broad to answer her specific question. She discovers that 28 of her 40 customers started buying directly from a local flour mill that offers loose atta at ₹32/kg versus her packed brand at ₹48/kg. She adjusts her order immediately and avoids stocking ₹12,000 worth of slow-moving inventory. The lesson: secondary data is your starting point, but when your decision is specific and time-sensitive, primary data collected with a simple, focused questionnaire can save you real money.
Common misconceptions to watch for
- Wrong belief: 'A bigger sample is always more accurate.' Correction: A randomly selected sample of 400 households beats a carelessly chosen sample of 4,000, because random selection removes bias — what matters is how you select, not just how many you select.
- Wrong belief: 'Secondary data is always the better choice because it is free and already available.' Correction: Secondary data may be outdated, collected for a different purpose, or too broad for your question; when your need is specific or time-sensitive, collecting primary data is often the right — and necessary — choice.
- Wrong belief: 'A census (100% survey) is the gold standard and sampling is just a compromise.' Correction: Censuses are expensive, slow, and only practical for certain situations; well-designed sample surveys like the NSO's National Sample Survey are standard professional tools that governments and researchers rely on precisely because they are faster and cheaper while still being statistically reliable.
Questions
A company wants customer satisfaction data across 50,000 stores. It can survey all 50,000 (₹50 lakhs, 6 months) or randomly select 500 (₹5 lakhs, 2 weeks). The launch deadline is 4 weeks. Which method is better and why?
- 1Identify the constraints: time (4-week deadline), budget (₹5-50 lakhs), and accuracy.A 6-month census exceeds the deadline. A 2-week sample fits. Time alone rules out the census, showing practicality matters as much as completeness.
Question 1 of 5 · easy
A survey of 1,000 randomly chosen Indian farmers vs. 200 randomly chosen farmers on crop yields. Both use the same random method. Which is true?
Quiz
Test yourself — pick an answer, then hit "Check" to see the explanation and your running score.
Question 1 of 5 · easy
A survey of 1,000 randomly chosen Indian farmers vs. 200 randomly chosen farmers on crop yields. Both use the same random method. Which is true?
Spotted an arithmetic error or unclear explanation? Suggest an edit — we fix things fast.