Organisation of Data
Before you can analyse any economic data — prices, wages, crop yields — you must first organise it. This chapter gives you the tools to turn a messy pile of raw numbers into a clean, meaningful frequency distribution.
Every economics exam question on statistics — from SCERT Kerala to CBSE — starts with organised data; mastering frequency distributions now also prepares you directly for the data interpretation questions that appear in board exams and for careers in accounting, commerce, or public policy.
Concept
Lots of students think…
"If an observation falls exactly on the boundary between two classes, it can go into either class — it does not matter which."
Actually…
Under the exclusive method used in economics, the upper boundary is excluded from each class. An observation of exactly ₹600 belongs to the class that starts at ₹600, not the one that ends at ₹600. There is no ambiguity once you apply the rule consistently.
By the end of this chapter, you will know how to take a messy pile of numbers — wages, prices, sales figures — and organise them into a clean table that tells a clear story. That one skill is the starting point for all of statistics.
Raw Data: The Unsorted Mess
Raw data is just a collection of numbers or facts recorded exactly as they were collected, in no particular order. On its own, raw data tells you almost nothing — you cannot spot a pattern just by staring at a random list. Organising data means sorting and grouping those observations so that patterns jump out immediately.
Imagine the Kerala Bureau of Statistics calls 200 fish-landing centres along the coast and asks each one: 'How many kilograms did you land today?' You now have 200 numbers scribbled in a notebook — 47 kg, 312 kg, 85 kg… in random order. Can you tell whether most centres land under 100 kg or over 500 kg? No. That is the problem raw data creates.
Qualitative vs Quantitative Data
Before you organise data, you need to know what kind it is. Qualitative data describes a quality or category — something you can name but not measure with a number, like the type of fish or the payment method used. Quantitative data is an actual number — like weight in kilograms or price in rupees — so you can add, subtract, and compare it.
At a wholesale fish market in Kozhikode, the type of fish (sardine, pomfret, tuna) is qualitative — you cannot add 'sardine + pomfret'. But the price per kilogram (₹120, ₹450, ₹380) is quantitative — you can average it or find the highest value. Two totally different types of data, needing different handling.
Discrete vs Continuous: Two Kinds of Numbers
Quantitative data splits into two sub-types. Discrete data can only be whole numbers — you cannot have 2.5 workers or 3.7 boats. Continuous data can take any value, including decimals — a catch weight could be 47.3 kg or a price could be ₹182.50. Knowing the difference matters when you decide how wide to make your groups.
At a prawn farm in Alappuzha, the number of workers employed each day is discrete — you hire 8 or 9 people, never 8.6. But the weight of the prawn harvest is continuous — today's batch might weigh 143.75 kg. Both are quantitative, but they behave differently.
Frequency Distribution: Grouping Numbers into Classes
A frequency distribution is a table that groups your observations into classes (ranges) and counts how many observations fall in each class — that count is called the frequency. Instead of listing every single number, you get a compact summary that immediately shows where most values cluster. Each range is called a class interval, and its size is the class width.
Meera tracks daily sales at her fish stall in Kozhikode for 60 days. Sales range from ₹8,000 to ₹72,000. She creates six classes — ₹0–12,000, ₹12,000–24,000, and so on — and counts how many days fall in each. She finds that 36 out of 60 days (60%) sit in the ₹24,000–48,000 band. Sixty random numbers told her nothing; this one table tells her her reliable income zone.
How Many Classes? The Square-Root Rule
Too few classes lump everything together and hide variation. Too many classes scatter the data into near-empty boxes. A simple starting point is the square-root rule: number of classes ≈ √N, where N is the total number of observations. It is a guideline, not a law — round up or down to make clean, convenient class widths.
A NABARD survey records the annual income of 100 small farmers in Palakkad district. The square-root rule suggests about √100 = 10 classes. If you use only 3 classes, you cannot see whether most farmers earn ₹80,000 or ₹2,00,000 — everything is blended. Ten classes give a clear picture of the income spread without drowning you in near-empty rows.
The Exclusive Method: No Overlap, No Confusion
When you list class intervals, you need a clear rule for boundary values — what happens when an observation falls exactly on the line between two classes? The exclusive method says: the upper limit of each class is excluded from that class and belongs to the next one. This means every single observation goes into exactly one class — no overlaps, no gaps.
A vegetable mandi in Thrissur records the daily wage of its workers. The classes are ₹500–600, ₹600–700, ₹700–800. Under the exclusive method, a worker earning exactly ₹600 goes into the ₹600–700 class, not the ₹500–600 class. Simple, unambiguous, and used as the standard rule in all your economics textbooks.
Notes
The full picture
Imagine the Kerala Bureau of Statistics has just finished a survey: 200 fish-landing centres, each reporting daily catch in kilograms. You now have 200 numbers scribbled in no particular order. Can you tell at a glance whether most landings are small or large? No — because the data is raw. Raw data is simply an unordered collection of observations. Organisation of data is the process of sorting and grouping those observations so patterns jump out. This is always the first step before any statistical analysis.
The first decision in organising data is identifying what kind of data you have. Qualitative data (also called attributes) describes characteristics that cannot be measured — the type of fish (sardine, tuna, pomfret), or the fishing method used (trawl, purse seine, hook-and-line). You cannot add or subtract these; you can only count how many observations fall in each category. Quantitative data, on the other hand, involves actual numbers — kilograms caught, price per kilogram in rupees, number of workers. Quantitative data divides further: discrete data takes only whole number values (3 boats, 12 workers), while continuous data can take any value in a range (47.3 kg, ₹182.50 per kg).
Once you know your data type, you build a frequency distribution. This is a table that groups observations into classes and records how many observations fall in each class — that count is the frequency. For example, take daily wages of 80 agricultural labourers in Thrissur district: instead of listing all 80 figures, you create groups like ₹400–500, ₹500–600, ₹600–700, ₹700–800, and count how many workers fall in each. You now have a compact picture of the wage structure at a glance. The range covered by each group is the class interval, and its width is the class width.
Choosing the right number of classes is a judgment call — a useful starting point is the square-root rule: number of classes ≈ √N, where N is the total number of observations. For 100 observations, try about 10 classes. Too few classes (say, 3) lump everything together and hide variation; too many (say, 50) scatter the data into near-empty boxes. Equal class widths are almost always preferred because they make comparison and later graphing straightforward. Unequal widths and open-ended classes (such as 'below ₹5,000' or 'above ₹50,000') are used only when data is very sparse at the extremes.
The exclusive method is the standard way to handle boundaries between classes in economics. In this method, the upper limit of each class is excluded — it belongs to the next class. If your classes are ₹500–600 and ₹600–700, an observation of exactly ₹600 goes into ₹600–700, not ₹500–600. This rule is simple and completely eliminates the problem of where to put boundary values. Every observation falls into exactly one class, with no overlaps and no gaps. Always use this method unless your textbook specifically instructs otherwise.
Constructing a frequency distribution table follows a clear sequence: (1) find the highest and lowest values to know your range; (2) decide how many classes you need using the square-root guideline; (3) calculate class width as range ÷ number of classes, then round to a convenient figure; (4) list the class intervals from lowest to highest; (5) tally each observation into its class; (6) write the frequency for each class. You can also add a relative frequency column showing each class's frequency as a percentage of total observations — this makes comparisons across datasets of different sizes fair and easy.
An Indian example
Meera runs a small wholesale fish market in Kozhikode. In one month she records the daily sale value (in rupees) for 60 trading days. The values range from ₹8,000 on slow monsoon days to ₹72,000 on peak festival days. Before she can understand her business, she groups these 60 figures into classes: ₹0–12,000 (4 days), ₹12,000–24,000 (9 days), ₹24,000–36,000 (17 days), ₹36,000–48,000 (19 days), ₹48,000–60,000 (8 days), ₹60,000–72,000 (3 days). Now the picture is clear: nearly 60% of days cluster in the ₹24,000–48,000 band — her reliable income zone. The few very high days (₹60,000+) are rare festival spikes, not the norm. She uses this organised table to decide when to hire extra staff and how much stock to order on average. Sixty raw numbers told her nothing; one frequency distribution told her everything she needed.
Common misconceptions to watch for
- Misconception: An observation that falls exactly on the boundary of two classes can go into either class. Correction: Under the exclusive method — which is standard in economics — the upper boundary is excluded from each class. So an observation of exactly ₹600 belongs to the class starting at ₹600, not the class ending at ₹600. There is no ambiguity.
- Misconception: You must always use exactly √N classes — it is a formula, not a guideline. Correction: The square-root rule is a starting point, not a law. For 64 observations it suggests 8 classes, but if your data is tightly clustered you might use 6, or if it is very spread out you might use 10. The goal is a table that clearly shows the shape of the data.
- Misconception: Qualitative data like crop type or fishing method needs to be converted to numbers (1 = rice, 2 = wheat, etc.) before you can organise it. Correction: No conversion is needed. You simply count how many observations belong to each named category and record those counts in a frequency table. Assigning arbitrary numbers to categories would imply a numerical relationship (is rice 'twice' wheat because 2 > 1?) that does not exist.
Questions
The Kerala Bureau of Statistics studies agricultural income across 75 smallholder farms in Thiruvananthapuram district. Income ranges from ₹18,000 to ₹1,92,000 annually. Organise this data into a frequency distribution table using equal class intervals (exclusive method). Should the number of classes be exactly 8?
- 1Apply the square-root rule to find the ideal number of classes
√N = √75 ≈ 8.66 Rounded: approximately 8–9 classes
The square-root rule is a guideline, not a law. With 75 observations, this suggests 8–9 classes. This helps avoid extremes: too many classes hide patterns, too few lose detail. We'll use 8 classes, but 7 or 9 would also be reasonable.
Question 1 of 5 · easy
In the exclusive method, an observation of exactly ₹50,000 income belongs in which class: ₹40,000–50,000 or ₹50,000–60,000?
Quiz
Test yourself — pick an answer, then hit "Check" to see the explanation and your running score.
Question 1 of 5 · easy
In the exclusive method, an observation of exactly ₹50,000 income belongs in which class: ₹40,000–50,000 or ₹50,000–60,000?
Spotted an arithmetic error or unclear explanation? Suggest an edit — we fix things fast.