Kerala HSE (SCERT) · Class 11 · Economics
Unit 1 · Chapter 7 · Statistics for Economics

Correlation

Correlation tells you whether two variables move together — and how strongly — giving you a powerful tool to spot patterns in economic data without falling into the trap of assuming one thing causes the other.

Correlation is tested in your board exam through both theory questions and numerical problems, and in real life it is how economists, businesses, and policymakers decide where to look for solutions — so mastering it puts you ahead in class and in any data-driven career including CA, B.Com, and Economics honours.

Concept

Quick myth-check

Lots of students think…

"A correlation of +1 proves that one variable is causing the other to change."

Actually…

r = +1 only means the two variables move together in a perfect positive linear pattern. A third variable could be driving both. Correlation shows a pattern — it never establishes cause.

By the end of this chapter you will understand what it means when two things move together — and, just as importantly, why that never proves one thing is causing the other.

What is Correlation?

Correlation tells you whether two variables tend to move together and in which direction. If both go up at the same time, that is a positive correlation. If one goes up while the other goes down, that is a negative correlation. If there is no pattern at all, the correlation is close to zero.

Real-life example

In Kerala, when the monsoon arrives, umbrella sales shoot up — and so do raincoat sales. Both rise together, so they have a positive correlation. Meanwhile, when petrol prices rise, the number of two-wheelers sold in Kochi tends to fall — one goes up, the other goes down, so that is a negative correlation.

The Scatter Diagram

A scatter diagram is the quickest way to see correlation. You plot one variable on the x-axis and the other on the y-axis, and put a dot for each pair of values. If the dots form a band that slopes upward, correlation is positive. If the band slopes downward, it is negative. If the dots are spread all over with no pattern, correlation is near zero.

Real-life example

Priya runs a stationery shop near a school in Thrissur. She plots her monthly notebook sales (x-axis) against her pen sales (y-axis) for twelve months. Every single dot lands in a tight band sloping upward — a clear visual sign that both items sell more at the same time of year.

Karl Pearson's Coefficient (r)

Karl Pearson gave us a formula that turns the scatter diagram into a single number called r. It always lands between -1 and +1. The formula compares how much the two variables vary together against how much each varies on its own — giving a clean, standard result you can use to compare any two datasets.

Real-life example

When Priya calculates Pearson's r for her notebook and pen sales, she gets r = 0.91. That single number confirms what her scatter diagram showed — a very strong positive relationship. She can now compare this to, say, the relationship between notebook sales and eraser sales (maybe r = 0.60 there) and know which pair is more closely linked.

Interpreting the Value of r

The closer r is to +1, the stronger and more positive the relationship. The closer to -1, the stronger and more negative. A value near 0 means almost no linear relationship. In practice you will see values like +0.82 (strong positive) or -0.45 (moderate negative) — rarely an exact 1, -1, or 0.

Real-life example

In India, research data often shows that states with higher per-capita income (x) also have higher literacy rates (y) — r might come out around +0.78, a strong positive. But the correlation between inflation rate and consumer confidence might be -0.55, a moderate negative — as prices keep rising, people feel less confident about spending.

Spearman's Rank Correlation

When your data is in the form of ranks — first, second, third — rather than measured numbers, you use Spearman's formula instead of Pearson's. Spearman looks at the difference between each item's two ranks and calculates r from that. Like Pearson, the result still falls between -1 and +1.

Real-life example

Ten students in a Commerce class in Kozhikode are ranked 1st to 10th by their teacher for public speaking ability, and then separately by another teacher for essay writing. If a student ranked 1st in speaking is also ranked 1st (or 2nd) in writing, the rankings agree and Spearman's r will be close to +1. If rankings are all jumbled — the best speaker turns out to be a poor writer — r will be much lower.

Correlation Does NOT Mean Causation

Just because two things move together does not mean one is causing the other. A hidden third factor can be pushing both at the same time. This is the most important idea in the whole chapter — and examiners love testing it.

Real-life example

Every summer in Chennai, ice cream sales rise. Reported cases of sunstroke also rise. The two are positively correlated — but eating ice cream does not cause sunstroke. The real cause of both is hot weather. A data-smart economist always asks: is there a third factor driving both variables? That habit separates careful analysis from jumping to wrong conclusions.

Notes

The shape of the scatter diagram tells you the direction and strength of correlation before you calculate a single number.

The full picture

Every time you notice that umbrella sales rise when it rains, or that petrol prices climb when crude oil gets expensive, you are already thinking about correlation. Correlation is a statistical measure that tells us how strongly two variables tend to move together and in which direction. If both variables increase together, the correlation is positive. If one rises while the other falls, the correlation is negative. And if there is no consistent pattern between them, the correlation is close to zero.

The most widely used measure is the Karl Pearson Correlation Coefficient, denoted by the letter r. It can only take values between −1 and +1. A value of +1 means a perfect positive linear relationship — every time one variable goes up, the other also goes up in a consistent linear fashion (not necessarily proportionally, just along a straight line with a positive slope). A value of −1 means a perfect negative linear relationship — one always rises as the other falls, following a straight line with a negative slope. A value of 0 means no linear relationship exists between the variables. In practice, r rarely hits exactly +1, −1, or 0; you will typically see values like +0.82 (strong positive) or −0.45 (moderate negative).

You can first get a visual feel for correlation by drawing a scatter diagram. Plot one variable on the x-axis and the other on the y-axis, then mark a dot for each data pair. If the dots form a tight band sloping upward, r is strongly positive. If they slope downward in a tight band, r is strongly negative. If the dots scatter randomly with no clear slope, r is near zero. After the visual check, you calculate r precisely using the Pearson formula — it compares how much the two variables vary together (their covariance) against how much each varies on its own (their standard deviations). This ratio always stays between −1 and +1, which is what makes it so convenient for comparing relationships across different datasets.

Spearman's Rank Correlation is a second method you need to know, particularly when your data is given as ranks rather than raw numbers. Suppose a panel of judges ranks ten students — 1st to 10th — by their public speaking ability, and separately ranks them by their essay scores. Here the data is ordinal (ranked order), not measured numerically. Pearson's formula is not designed for this type of data, so you use Spearman's formula instead, which is based on the difference between each student's two ranks. If the two rankings agree closely, Spearman's r is close to +1; if they are very different, it falls toward −1.

The single most important idea to take away from this chapter is that correlation does not imply causation. Two variables can move together for a reason completely unrelated to either of them. For example, ice cream sales and cases of sunstroke both rise in the summer months — they are positively correlated — but eating ice cream does not cause sunstroke. A third variable, hot weather, causes both. Examiners at the Kerala HSE board test this idea directly, so always check: is there a plausible third factor driving both variables? That question is the hallmark of careful statistical thinking.

An Indian example

Priya runs a small stationery and notebook shop near a government school in Thrissur. She notices that every time a new academic year starts, both her notebook sales and her pen sales shoot up together. Curious, she records monthly sales for twelve months: notebooks (in packs) and pens (in dozens). When she plots a scatter diagram, all the dots cluster tightly along an upward slope. She calculates the Karl Pearson Correlation Coefficient and gets r = 0.91 — a very strong positive relationship. This tells her that planning her stock for both items together makes sense: if she expects notebook sales to be unusually high in a month, pen sales will very likely be high too. However, Priya is careful not to say that selling notebooks causes pen sales to rise — both are driven by the same underlying factor, the school calendar. She uses the correlation purely to plan her inventory and avoid running out of stock during rush periods.

Common misconceptions to watch for

  • Wrong belief: 'A correlation of +1 proves that one variable causes the other.' This is false. r = +1 only means the two variables move together in a perfect positive linear fashion. A third variable could be causing both — correlation shows a pattern, never a cause.
  • Wrong belief: 'If correlation is zero, the two variables have nothing to do with each other.' This is false. Zero correlation means no linear relationship, but a strong non-linear (curved) relationship can still exist. For example, human physical strength rises from childhood to about age 25–30, then declines — an inverted-U-shaped curve (like a hill) — but a linear correlation of those two variables could be close to zero.
  • Wrong belief: 'Pearson and Spearman correlation always give the same answer, so it does not matter which one you use.' This is false. Pearson works on raw numerical data and is affected by outliers and skewed distributions. Spearman works on ranks and is more reliable when data is ordinal or contains extreme values. For rank-based data you must use Spearman; using Pearson on ranks gives a misleading result.

Questions

Worked example

An economist studying Kerala's coffee production analysed 8 years of data (2016–2023). She recorded annual rainfall days and coffee production. Data pairs: (65, 120), (72, 135), (58, 110), (80, 145), (70, 125), (78, 140), (62, 115), (68, 130), where X = rainfall days and Y = production (1000 kg). Calculate the Karl Pearson Correlation Coefficient.

1 / 5
  1. 1
    Organise data: X = rainfall days, Y = coffee production (1000 kg).
    Naming variables clarifies the order for applying the correlation formula. This distinguishes the independent variable (rainfall) from the outcome (production).
Reveal one step at a time. Read each before the next.
Practice

Question 1 of 5 · easy

0 / 0 correct

Ice cream vendors and swimming pool visits both increase during summer, showing r = +0.92. Which correctly interprets this?

Quiz

Test yourself — pick an answer, then hit "Check" to see the explanation and your running score.

Quiz

Question 1 of 5 · easy

0 / 5 correct

Ice cream vendors and swimming pool visits both increase during summer, showing r = +0.92. Which correctly interprets this?

How sure are you?
Answer to see your score.

Spotted an arithmetic error or unclear explanation? Suggest an edit — we fix things fast.

Index Numbers →