Three missions of checking other people’s numbers. Now the numbers are yours — which is harder, because the person most likely to fool you with your own data is you.
🧪Step 1 · Define It Before You Measure It
A variable is anything that changes and can be recorded. The trick is picking ones defined so tightly that you’d write down the same number twice if you measured twice.
🌫️Too fuzzy to measure“How much I read.” Minutes or pages? Does a comic count? Does the reading for school count? By Thursday you’ll be quietly changing the rules on Monday.
🎯Tight enough to trust“Minutes I spent reading, timed on a clock, not counting schoolwork.” Even a feeling works if you define it: “how rested I felt this morning, 1–5, where 3 means normal.”
✉️
Write your hypothesis before day one
A hypothesis is a specific guess, in one sentence, written down before any data exists — like sealing a prediction in an envelope. That feels like a formality, and it isn’t: predicting in advance is exactly what stops you from cherry-picking a pattern out of the noise afterward.
one sentencewritten firstspecific enough to be wrong
📓Step 2 · Keep an Honest Data Log
Seven rows, one per day, three columns. Record at the same time each day — sleep in the morning, reading before bed. Here’s one real-shaped week.
Two rules that protect the whole project. First: never fill in the week on Sunday from memory — memory is not a measuring instrument. It rounds, it flatters, and it bends numbers toward whatever you already believe. Second: if you miss a day, write “missing.” Don’t guess and don’t skip it silently. A blank you can see is honest data. An invented number never is.
🧮Step 3 · Do the Math You Already Know
Mission one is about to happen in your own bed. Sort each column, then run the numbers.
Two long weekend sleeps pulled the mean up. The nights add to 7 + 7 + 6 + 7 + 9 + 10 + 10 = 56, and 56 ÷ 7 = 8 hours. But sorted, the middle value is 7 — so a typical night is an hour shorter than the mean says. Meanwhile 175 minutes of reading is 2 hours and 55 minutes for the week, and 175 ÷ 7 = 25 minutes a day, which happens to be the median too.
📊Step 4 · Chart It Fairly
First a bar chart of reading minutes by day — bars starting at zero, because you know exactly what happens when they don’t.
Then a scatter plot: sleep along the bottom, reading up the side, one dot per day — seven dots, because you have seven days.
And if the dots look like spilled rice? That’s a genuine result too. “No pattern” is an answer, not a failure — and reporting it honestly is exactly the habit that makes your other findings believable.
🔗Step 5 · A Correlation Is a Clue, Not a Verdict
This is where almost everyone online goes wrong — and where you won’t. Whenever two things move together, there are always four explanations on the table.
🎈Look again at that week. The two highest sleep days and the two highest reading days are Saturday and Sunday — which are also the only two days with no school. Free time could easily be raising both numbers on its own. That’s a hidden variable, and your scatter plot has no way to see it.
👟Bigger shoes, better readingKids with larger shoe sizes really do tend to read better. The hidden variable is age: older kids have bigger feet and more years of reading practice. Age is doing all the work.
🍦Ice cream and swimmingIce cream sales and the number of people swimming rise and fall together all summer. Neither causes the other — hot weather causes both.
⚖️
So how does anyone ever prove causation? A fair test.
Change one thing on purpose and hold everything else steady. Pick two similar weeks, and in the second one deliberately go to bed 45 minutes earlier while keeping the rest of your schedule the same. Now a difference has somewhere to come from.
change ONE thinghold the rest steadydecide the plan first
✍️Step 6 · Write a Conclusion You’d Defend
Now the hardest sentence in the whole subject — the honest one. It says exactly as much as your data supports, and not one word more.
Name your own boundary too. Your log holds 7 data points from exactly one person. It describes your week and says nothing about kids in general. That isn’t a weakness — every real study has boundaries, and stating yours out loud is what makes the rest of your findings believable.
🔑Key Terms
📐VariableAnything that changes and can be measured or recorded — hours of sleep, minutes read, a mood score from 1 to 5.
✉️HypothesisA specific, testable guess written down before you collect data, so you can’t quietly change what you were looking for.
📓Data logA record where you write each measurement at the time you take it, instead of reconstructing it later from memory.
🔵Scatter plotA chart where each dot is one observation, placed by its value on two variables at once.
🔗CorrelationA pattern where two variables tend to move together — both rising, or one rising as the other falls.
⚡CausationWhen one thing actually makes another happen. A much stronger claim than correlation, and much harder to prove.
🎭Hidden variableA third factor you didn’t measure that is quietly causing both of the things you did measure.
🎲Spurious correlationTwo variables whose graphs match closely by pure coincidence, with no real connection at all.
Two more worth knowing: a trend is the overall direction your points move even when individual days bounce around — and a conclusion is the statement of what your data does and does not show. Honest conclusions come with their own fence built in: “in my one week of data…” marks exactly where the evidence stops.
🌍Where You’ll See This in Real Life
🩺Doctors’ officesDoctors regularly ask patients to keep a sleep diary or a headache log for a couple of weeks before an appointment — precisely because memory is an unreliable measuring instrument. A written log recorded day by day reveals patterns nobody can reconstruct accurately from recollection.
🏃Sports teams and coachesCoaching staff log training minutes, sleep, and soreness scores, then look for patterns across weeks. Good ones treat a correlation as a lead worth investigating rather than a proven cause — and they test changes one at a time so they can tell what actually made the difference.
🏛️“Data” is a plural word borrowed from Latin. A single piece of data is technically a datum, from a Latin word meaning “something given” — which is a lovely way to think about a measurement: something the world hands you, whether or not it’s what you hoped for.
📅One week of daily tracking gives you 7 data points. A full year gives you 365. That’s why scientists repeat measurements instead of trusting a single reading — more points make a real pattern much easier to separate from random bounce.
📌Remember This
1Good data starts before the data does: define your variables tightly, write your hypothesis first, and log every day at the same time instead of reconstructing from memory.
2A correlation always has four possible explanations — A causes B, B causes A, a hidden variable causes both, or coincidence — so it’s a clue to investigate, never a verdict.
3The most valuable sentence you can write is an honest conclusion that names its own limits: what you found, what you can’t tell yet, and what you’d measure next.
🤔 Think about it
Apps and fitness trackers collect this kind of data about people automatically, all day, without anyone writing anything down. What do you gain from that — and what do you give up?
If you ran your week and found no pattern at all, would you be tempted not to share it? Why do you think “we found nothing” results are so rarely reported — and what problems might that cause?
⭐Remember: when you make the chart, you choose where the axis starts, which week to show, and which average to print. Every trick you learned to catch is now a temptation you get to refuse — and refusing it is exactly why people will believe your numbers.
✏️ ClickClass Anchor Chart · Your Data Story: Collect It, Chart It, Tell the Truth About It