Topline

Every health number carries a population of origin, an error range and a threshold drawn by committee. Here is how to read one without overreacting to it.

A health number is not a fact about your body. It is the output of a formula that was fitted to some population, using measurements that carry error, compared against a threshold that a committee chose. All four of those things can be wrong for you at once, and none of them are visible in the number itself.

This is the most useful thing we can teach, and it argues against taking our own calculators too literally. A person who understands where an estimate comes from will get more out of it than a person who treats it as a reading off an instrument, because they will know which changes mean something and which are noise. What follows is the set of questions worth asking of any health number, whether it comes from this site, an app, a smartwatch or a clinic.

Ask which population it came from

Every formula was derived from a specific group of people, and its accuracy degrades as you differ from that group.

The ratio behind body mass index was published by Adolphe Quetelet in the 1830s to describe how traits distribute around a population mean. He was a statistician interested in the average man, and he made no claim about individual health. The thresholds layered onto it much later came largely from European-ancestry cohorts, which is why the WHO Expert Consultation reported in the Lancet (363:157-163, 2004) that people of Asian descent carry higher body fat and greater cardiometabolic risk at the same index value, and recommended additional public health action points at roughly 23.0 and 27.5. Our BMI calculator shows both sets of thresholds for that reason.

The pattern repeats across every tool here. The circumference equations behind our body fat calculator come from Hodgdon and Beckett's work for the US Navy (Naval Health Research Center Report 84-29, 1984), developed on military personnel who were younger, fitter and more uniformly built than the general population. The Mifflin-St Jeor equation, published by Mifflin and colleagues in the American Journal of Clinical Nutrition (51:241-247, 1990), was fitted to a few hundred adults across a range of body sizes, and remains the best general predictor of resting metabolism precisely because that sample was reasonably representative. Age-predicted maximum heart rate, from Tanaka, Monahan and Seals in the Journal of the American College of Cardiology (37:153-156, 2001), was built from a meta-analysis of healthy adults.

The practical question is not whether a formula is good but whether you resemble the people it was built on. A seventy-year-old woman using an equation validated on young men is not getting a bad estimate through any fault of the equation.

Ask how much error it carries

Almost no health number is reported with its error range attached, which is the single largest cause of over-reading.

The ranges are not small. Predictive energy equations sit within roughly ten percent of measured expenditure for most adults, and further out for a meaningful minority. Run a forty-year-old man of 82 kg and 178 cm through the TDEE calculator at a moderate activity level and Mifflin-St Jeor returns about 2,693 kcal. A ten percent band around that spans 2,424 to 2,962 kcal. That is a spread of more than five hundred calories, which is larger than most people's deliberate deficit.

Metric Typical error against a reference method
Predicted energy expenditure around ±10%, wider for some individuals
Tape-based body fat roughly 3 to 4 percentage points
Age-predicted maximum heart rate roughly ±10 to 12 beats per minute
Body mass index as a proxy for fatness varies with muscle mass and fat distribution

Each of those has a consequence. Three to four percentage points of error on body fat means a reading of 21% is compatible with a true value in the high teens or the middle twenties, so a change from 21% to 19% between two measurements may be entirely measurement. Ten to twelve beats per minute on maximum heart rate is enough to move every training zone: for a forty-five-year-old, the heart rate zone calculator returns a maximum of 177 and an aerobic zone of 106 to 124 beats, but if that person's true maximum is 189 the honest zone is 113 to 132, and if it is 165 the zone is 99 to 115. Training by the number rather than by breathing and perceived effort will put a substantial minority of people in the wrong zone entirely.

Error is not a reason to discard a number. It is a reason to stop reading its last digit.

Ask what the threshold actually represents

Category boundaries feel like biology and are almost always administrative. They exist because someone had to draw a line for reporting, screening or reimbursement, and a line has to go somewhere.

Watch what happens at one. At 178 cm, our BMI calculator returns 24.9 and the label healthy weight at 79 kg, and 25.2 with the label overweight at 80 kg. One kilogram, less than a day's variation in fluid and gut contents, moves a person across a category boundary. Nothing about their cardiovascular risk changed overnight. The WHO Technical Report Series 894 (2000), which codified those bands, presents them as pragmatic cut-offs for population comparison rather than as points where risk jumps.

Risk relationships in the underlying data are almost always continuous and gently curved. Thresholds impose steps onto curves, which is useful for counting populations and misleading for describing individuals. The correct reading of a value near any boundary is that you are near a boundary, which carries almost no information.

Ask whether it predicts or explains

Most health metrics are associations observed across populations. An association tells you how outcomes differ, on average, between groups that differ on some measure. It does not tell you that the measure caused the difference, and it does not tell you what will happen to one person who changes it.

Two failure modes follow. Confounding is the familiar one: waist circumference predicts cardiovascular risk, and it also travels with diet, activity, sleep, income and a dozen other things that independently affect risk. Reverse causation is the underrated one, because it looks identical in the data. Lower grip strength predicts mortality, but early illness reduces grip strength, so some of that relationship runs backwards from the direction people assume.

There is also the question of how much to trust any single published finding. Ioannidis made the general case in PLoS Medicine (2005) that a large share of published research findings are false, driven by small studies, small effects, flexible analysis and the incentive to publish something novel. The practical consequence is not cynicism but patience: a finding replicated across large cohorts on several continents, as the body mass index and mortality relationship has been, deserves considerably more weight than a single striking result.

This is why we prefer to say what a number is associated with rather than what it does to you, and it is set out further in our editorial policy.

Ask whether you are looking at a point or a trend

This is the most immediately useful habit on the page, and the cheapest to adopt.

A single measurement is dominated by noise. Body weight moves by one to two kilograms within a day on fluid, glycogen, sodium and gut contents, none of it adipose tissue. Each gram of stored glycogen carries roughly three grams of water, so a high-carbohydrate weekend or a hard training session can add weight that has nothing to do with fat. Blood pressure varies with posture, time of day, caffeine and the experience of having it measured. Resting heart rate rises with poor sleep and dehydration.

Against that, the trend is comparatively clean. Random error is as likely to sit above the truth as below it, so it partly cancels across repeated measurements while a real change accumulates. This is the entire reason a seven-day rolling weight average is more informative than any morning's figure, and why a tape measurement repeated monthly by the same person at the same landmark beats a one-off body composition scan for tracking purposes.

Three rules make this operational. Measure under standardised conditions, because consistency of method matters more than accuracy of method when you are looking for change. Compare periods rather than points, week against week rather than day against day. And decide in advance how much change would be meaningful, which for most metrics means a movement larger than the measurement error described above.

Ask what the number is for

A metric is fit for a purpose, and most misreadings come from using one outside its purpose.

Screening tools are designed to sort a population into people worth examining further and people who probably are not. They are deliberately tuned to over-refer, because in screening a false alarm costs a follow-up appointment while a miss costs a diagnosis. Body mass index is a screening tool in exactly this sense. Reading a screening result as a diagnosis is the most common error in consumer health, and it runs in both directions: a flag is not a disease, and the absence of a flag is not clearance.

Estimates are a different category. A predicted energy requirement is a starting hypothesis to be calibrated against what actually happens to your weight over three or four weeks. It is the beginning of a measurement process, not the result of one.

Windows are a third. A due date from the pregnancy due date calculator is a single date standing in for a distribution: it comes from Naegele's rule, adding 280 days to the last menstrual period, and most births occur across a spread of weeks around it rather than on it. Reporting a distribution as a date is a convention of convenience, and gestational age explained covers how wide the real window is.

Reference ranges are a fourth, and the least intuitive. A laboratory reference range is usually the central 95% of a healthy reference population, which means one healthy person in twenty falls outside it by construction. Order enough independent tests and a result outside the range becomes likely for reasons that have nothing to do with disease.

Putting it together

None of this means the numbers are worthless. It means each one carries a set of qualifiers that should travel with it, and a reader who keeps them attached will be right more often.

Held that way, a health number is a prompt to look more carefully at something, and occasionally a prompt to see a clinician. What it is not is a verdict. The reader who improves fastest is usually the one who checks fewest numbers, measures them consistently, and reads them over months.

Two habits carry most of the benefit. Read any metric alongside a second one that fails differently: body mass index next to a waist measurement, as waist-to-height ratio explained sets out, or scale weight next to a tape figure, as in weight loss versus fat loss. And prefer the direction of travel to the absolute value, because the direction is more robust to every error described above. A body mass index that has climbed four units in five years is telling you something the same value held steady for five years is not.

It is also worth noticing when measurement stops being useful. If checking a number has become something you do several times a day, if a reading determines your mood, or if the pursuit of a target has begun to narrow how you eat, train or live, the problem is no longer accuracy and a better estimate will not fix it. That is worth raising with a clinician or a psychological professional. Anyone whose relationship with food, exercise or body weight has become a source of distress deserves proper support rather than a more precise instrument.

Our full set of estimators is on the calculators page, and the honest summary of all of them is this: they turn measurements you can take at home into figures that are usually about right, for a person roughly like the ones each formula was built on, with an error range you should keep in mind. That is genuinely useful. It is also considerably less than most health numbers are asked to carry.