CBSE · Class 11 · Economics
Unit 1 · Chapter 7 · Statistics for Economics

Correlation

Correlation tells you whether two economic variables move together and how strongly — mastering Karl Pearson's coefficient, Spearman's rank method, and scatter diagrams gives you a powerful tool to find real patterns in data rather than mistaking coincidence for connection.

Correlation is tested directly in CBSE board exams through calculation problems and conceptual questions, and it is the foundation for every data-driven career — whether you go into economics, business research, CA, or banking, you will use it to separate real patterns from coincidence.

Concept

Quick myth-check

Lots of students think…

"If two things go up and down together and the r value is high, one must be causing the other."

Actually…

Correlation only tells you that two variables move together — it says nothing about cause. Ice cream sales and heatstroke cases both rise every summer because of hot weather, not because ice cream causes heatstroke. Always look for a third hidden factor before claiming any cause.

By the end of this chapter you will know how to find out whether two things — like temperature and cold-drink sales — really move together, and how to measure exactly how strong that link is. You will never again confuse a coincidence with a real pattern.

What is Correlation?

Correlation tells you whether two variables move together — and if so, how strongly. It does NOT tell you that one thing causes the other; it only shows a pattern. Spotting that difference is the most important habit in statistics.

Real-life example

In Coimbatore, Priya notices that every month her kirana store sells more cold drinks when the temperature is higher. Sales and temperature seem to move together. That is correlation — but the heat is not 'caused' by the cold drinks; they are just linked by the season.

Scatter Diagrams — See the Pattern First

A scatter diagram is a dot graph. You put one variable on the X-axis and the other on the Y-axis, then plot one dot per observation. The shape of the dots tells you the story before you calculate anything — always draw this first.

Real-life example

Priya records temperature and sales for 8 months and plots each month as a dot. The dots slope upward from left to right — like a hill going up — so she can already see that higher temperature goes with higher sales, even before doing any maths.

Positive and Negative Correlation

Positive correlation means both variables go up together (or both go down). Negative correlation means when one goes up, the other goes down. You can spot both instantly on a scatter diagram before calculating anything.

Real-life example

Positive: as a farmer's land (in acres) increases, his rice output (in kg) also increases — both go up together. Negative: as the price of onions rises at Delhi's Azadpur Mandi, the quantity people buy falls — one goes up, the other goes down.

Karl Pearson's r — Putting a Number to It

Karl Pearson's coefficient, written as r, is a single number between −1 and +1 that tells you how strong and in which direction the linear relationship is. Divide how much the two variables vary together (covariance) by the product of their individual spreads (standard deviations). Use this method when both variables are measured on a proper scale — like rupees or kilograms.

Real-life example

Priya calculates Pearson's r for her temperature-vs-sales data and gets r = 0.91. That is very close to +1, so she now knows the positive link between temperature and cold-drink sales is strong and consistent — not just a coincidence she noticed in her head.

Spearman's Rank Correlation — When You Have Rankings

Spearman's method is for ranked data — like 1st, 2nd, 3rd positions — or when one very extreme value would mess up Pearson's result. You replace each value with its rank and use the differences in ranks (D) to calculate the coefficient. It gives you the same −1 to +1 range.

Real-life example

A teacher in Kochi ranks 10 students by their Maths score and by their Science score. She cannot use Pearson's r because these are positions, not measurements. She uses Spearman's formula with the rank differences and gets r = 0.85 — showing students who do well in Maths tend to do well in Science too.

Correlation is NOT Causation

Just because two things move together does not mean one causes the other. There is almost always a third factor — or it could be pure coincidence. Always ask: 'Is there something else driving both of these?' before drawing any conclusion.

Real-life example

Every summer in India, both ice cream sales and the number of heatstroke cases rise together. Pearson's r for these two would be very high — but ice cream does not cause heatstroke. Hot weather is the real driver of both. If a government wrongly banned ice cream to cut heatstroke, it would change nothing. This is why the distinction matters.

Notes

Reading a scatter diagram is the first step in any correlation analysis — the shape of the point cloud tells you the direction and rough strength before you calculate r.

The full picture

When you look at economic data, you often wonder: do two variables move together? Does more rainfall mean more rice output? Does higher household income lead to more spending? Correlation is the statistical tool that answers this question. It measures the strength and direction of a relationship between two variables. Notice that correlation only describes a pattern — it does not tell you that one variable causes the other. For example, both umbrella sales and cold and flu cases rise in the monsoon season, but umbrellas do not cause illness. Keeping the correlation-versus-causation distinction clear is the single most important habit in statistics.

The first tool you will use is the scatter diagram. Plot one variable on the X-axis and the other on the Y-axis, one point per observation. The shape that emerges tells you everything before you calculate a single number. When points slope upward from left to right, you have positive correlation — both variables rise together. When points slope downward, you have negative correlation — one rises as the other falls. When points are scattered randomly with no pattern, there is no (or very weak) correlation. Always draw a scatter diagram first; it catches non-linear curves and stray outliers that any single number can hide.

Karl Pearson's correlation coefficient, written as r, gives you a precise number. It ranges from −1 to +1. A value of +1 means a perfect positive linear relationship; −1 means a perfect negative linear relationship; 0 means no linear relationship. The formula divides the covariance (how much the two variables vary together) by the product of their standard deviations. In practice, r above 0.7 (or below −0.7) is considered strong, between 0.3 and 0.7 (or −0.3 and −0.7) is moderate, and below 0.3 is weak. Pearson's method requires both variables to be measured on a continuous scale — such as income in rupees or weight in kilograms — and the relationship should be roughly linear, which is exactly why you draw the scatter diagram first.

Spearman's rank correlation coefficient solves a real problem. When your data are ordinal — ranked positions rather than precise measurements — or when a few extreme outliers are pulling Pearson's r away from the true picture, Spearman is the better choice. You convert each variable's raw values into ranks (1st, 2nd, 3rd…) and then apply a simple formula using the difference in ranks (written as D). Because Spearman works with positions rather than actual values, an extreme outlier occupies the same rank boundary (first or last) no matter how extreme its measured value is, leaving the overall coefficient stable. This makes Spearman ideal for comparing students by exam rank and attendance rank, or judging cities by quality-of-life score.

Here is how both methods work together in practice. Suppose a school teacher in Kochi records the height in centimetres and weight in kilograms of 10 students. She first plots a scatter diagram — the points slope upward, suggesting positive correlation. She calculates Pearson's r = 0.88, confirming strong positive linear correlation: taller students tend to weigh more. Now she also ranks students by height and by weight separately, then applies Spearman's formula. She gets r_s = 0.91, slightly higher because one very short but heavy student (an outlier) was dragging Pearson's r down slightly. Both coefficients together paint a reliable picture. She presents the scatter diagram alongside both numbers in her school project — this is exactly the approach NCERT expects.

An Indian example

Priya runs a small kirana store in Coimbatore. She notices her monthly cold-drink sales (in ₹ thousands) seem to rise and fall with local temperature. She records data for 8 months: in April (42°C) she sold ₹38,000; in June (28°C, monsoon) she sold ₹18,000; in December (20°C) she sold ₹12,000. She plots these on a scatter diagram — the points slope clearly upward from left to right, so she expects positive correlation. She calculates Pearson's r and gets 0.91 — strong positive correlation. Her supplier then offers a bulk discount if she orders in advance. Because of her correlation analysis, Priya confidently orders extra stock before March each year, cutting her per-unit cost by 8%. She does not assume temperature causes her sales to rise; she knows both are driven by the season. But the pattern is consistent enough to plan inventory around — and that is exactly what correlation is for.

Key concepts covered

  • Karl Pearson
  • Spearman's rank
  • Scatter diagrams

Common misconceptions to watch for

  • Misconception: 'If r is high, one variable must be causing the other to change.' Correction: Correlation only shows that two variables move together — it never proves causation. In India, both ice cream sales and heatstroke cases rise together every summer, but ice cream does not cause heatstroke; a common third factor (hot weather) drives both. Always ask whether a third variable could explain the pattern before drawing any causal conclusion.
  • Misconception: 'Pearson's r works for any type of data, including ranked lists.' Correction: Pearson's r requires continuous, interval-or-ratio-scale data (like income in ₹ or height in cm) and a linear relationship. If you are comparing students' rank in Maths with their rank in Science — ordinal data — you should use Spearman's rank correlation. Using Pearson on rank numbers treats a difference from rank 1 to rank 2 as the same as a difference from rank 9 to rank 10, which is not meaningful in a ranked list.
  • Misconception: 'An r value near zero means the two variables are completely unrelated.' Correction: Pearson's r only measures linear (straight-line) relationships. Two variables can have r ≈ 0 and still be strongly related through a curve. For example, electricity demand is high in both extreme summer (air conditioners) and extreme winter (heaters) but low in mild weather — a U-shape on a scatter diagram — yet Pearson's r for this pattern comes out near zero. Always draw the scatter diagram first to check the shape before trusting any single coefficient.

Questions

Worked example

A Bangalore mall analyst examines whether customer footfall (daily visitors) correlates with ice cream sales (₹ thousands). Data from 6 days: Footfall: 1200, 1500, 1800, 2100, 2400, 2700. Sales: 25, 32, 38, 45, 51, 58. Calculate Karl Pearson's r.

1 / 6
  1. 1
    Organise data into columns X (footfall) and Y (ice cream sales in ₹ thousands).
    X: 1200, 1500, 1800, 2100, 2400, 2700
    Y: 25, 32, 38, 45, 51, 58
    Converting sales to thousands simplifies calculation. Clear organisation prevents errors in deviations and sums.
Reveal one step at a time. Read each before the next.
Practice

Question 1 of 5 · easy

0 / 0 correct

House prices in Delhi rose 25% over five years; salaries also rose 25%. A realtor claims rising salaries caused house prices to rise. Why may this conclusion be incorrect?

Quiz

Test yourself — pick an answer, then hit "Check" to see the explanation and your running score.

Quiz

Question 1 of 5 · easy

0 / 5 correct

House prices in Delhi rose 25% over five years; salaries also rose 25%. A realtor claims rising salaries caused house prices to rise. Why may this conclusion be incorrect?

How sure are you?
Answer to see your score.

Spotted an arithmetic error or unclear explanation? Suggest an edit — we fix things fast.

Index Numbers →