Topline

Tracked intake is typically wrong by ten to thirty percent, and the errors run one way. Here is where they come from and why tracking is still worth doing.

A tracked calorie total is usually wrong by ten to thirty percent, and the error is not random. It runs in one direction: people record less than they eat. The magnitude varies enormously between individuals, which is the part that makes it hard to correct for, and it is largely unrelated to how conscientious the person is being.

None of that makes tracking pointless. It makes it a tool with a known bias, and a biased instrument used consistently is still useful. It just cannot be used the way most people use it. The error stacks up from at least six separate sources, each measurable, and knowing which ones dominate tells you where to spend your effort.

Our companion piece on reading nutrition labels covers the regulatory side of this: what label numbers legally mean and the tolerances manufacturers work within. This article is about the errors that happen on your side of the packet.

The size of the self-report problem

The definitive work here uses doubly labelled water, a method that measures total energy expenditure by tracking the disappearance rates of two stable isotopes. In a weight-stable person, expenditure equals intake, which gives researchers an objective yardstick to hold against what participants reported eating.

Lichtman and colleagues, in the New England Journal of Medicine (327:1893-1898, 1992), applied this to adults who reported being unable to lose weight on intakes they described as low. Reported intake was under-recorded by close to half, and physical activity was over-reported by around a quarter. The participants were not lying; several were distressed by the results and had genuinely believed their own records.

Schoeller reviewed the broader literature in Metabolism (1995) and found under-reporting to be the rule rather than the exception across study populations, with the magnitude varying by individual and tending to be larger in people with higher body weight and in women. The general finding has been reproduced many times since across different populations and methods.

Two features matter for anyone using an app. The bias is directional, so it does not cancel out over a week. And it is person-specific, so a population average correction factor is not much help: you might be the person under-recording by five percent or the one under-recording by thirty-five.

Portion estimation, which is harder than it looks

The largest tractable error is the difference between the amount you think you ate and the amount you ate.

Human volume estimation is poor and it is biased. People underestimate large portions more than small ones, so the error grows with the size of the meal. Calorie-dense items produce the largest absolute errors because a small volume misjudgement carries a lot of energy: a tablespoon of olive oil is about 120 kcal, and the difference between a level and a generous tablespoon is real money. Peanut butter, cheese, cream, nuts, granola, dressings and anything cooked in fat all behave the same way.

The items most often omitted entirely are worse. Cooking oil left in the pan, the splash of milk in four coffees, sauces, the handful of something while cooking, alcohol, the last third of a child's meal. Individually these are trivial. Together they routinely account for 200 to 400 kcal a day, and none of them feel like eating.

The fix is unglamorous and effective: weigh the dense things. A kitchen scale removes almost all of this error for the foods where it matters most, and vegetables can be estimated indefinitely without consequence. Weighing cooked rather than raw introduces its own error, since water content varies with cooking method, so pick one convention and stay with it.

The data you are logging against

Even a perfectly weighed portion is only as good as the entry you attach it to, and two sources of that data behave badly in different ways.

Crowd-sourced database entries

Most tracking apps draw on crowd-sourced databases, and the entries are inconsistent in ways the interface conceals. Search for a common food and you will typically find a dozen entries with different energy values, most of them user-submitted, none of them verified.

Several failure modes are common. Entries for the same food differ by twenty percent or more with no way to tell which is right. Serving sizes are attached to ambiguous units, so a "medium" anything is whatever the submitter meant. Recipes are entered as single items with no visibility into their composition. And branded entries persist after the manufacturer has reformulated the product.

The mitigation is to build a small, stable set of entries you trust (verified national database entries or manufacturer figures where available) and reuse them. The absolute accuracy improves somewhat. More importantly, the consistency improves a lot, and consistency is what makes the trend readable.

Restaurant and prepared food

Eating out breaks tracking more than any other single behaviour, because you control neither the portion nor the preparation.

Urban and colleagues measured the energy content of restaurant meals against their stated values and reported the results in the Journal of the American Dietetic Association (2010). Measured energy exceeded stated energy on average, with substantial variation between individual items and between repeat purchases of the same dish. Portion inconsistency between one visit and another is a large part of this, along with cooking fat that no menu figure captures.

The practical position is that a restaurant meal is an estimate with wide error bars, and no amount of care in the app narrows them. Recording a plausible figure and accepting the uncertainty beats either omitting the meal or spending twenty minutes constructing a false precision.

The Atwater factors are averages

Even a perfectly weighed food logged against a perfectly accurate database entry carries error, because the energy value itself is a calculation rather than a measurement.

Calorie figures come from Atwater factors: 4 kcal per gram for protein and carbohydrate, 9 for fat, 7 for alcohol. Wilbur Atwater derived these in the late nineteenth century by burning foods in a calorimeter and subtracting estimated losses in urine and faeces. They are category averages, and they have known biases.

Novotny, Gebauer and Baer, in the American Journal of Clinical Nutrition (2012), measured metabolisable energy from almonds and found roughly a fifth fewer calories available than the Atwater calculation predicts, because intact cell walls resist digestion. Similar overestimates have been reported for other whole nuts and for some high-fibre foods.

The bias runs the other way for processed food. Carmody, Weintraub and Wrangham demonstrated in the Proceedings of the National Academy of Sciences (2011) that cooking and mechanical processing increase the energy actually extracted from food, by disrupting cell structure and gelatinising starch. A label does not distinguish between raw and cooked forms of the same nominal ingredient.

So the systematic effect is that labels slightly overstate the energy you extract from whole, fibrous, minimally processed food, and understate it at the highly processed end. Fibre intake and gut microbiota composition modulate this further, by amounts that are real but not currently quantifiable for an individual.

Adding the errors up

Take a plausible day for a 35-year-old man, 178 cm, 75 kg, moderately active. His maintenance estimate from the TDEE calculator is 2,623 kcal, and the steady fat-loss preset gives him 2,230 kcal, a nominal deficit of about 393 kcal.

Now stack modest errors. Label tolerances permit measured energy up to twenty percent above declared for calories, and packaged items are a real share of his day. Two unweighed dense items add perhaps 80 kcal. Cooking oil not fully counted adds 60. One restaurant meal in the week runs 150 above its stated figure. Database entries drift a few percent in the generous direction.

None of these is egregious. Together they can easily reach 300 to 400 kcal a day, which is his entire deficit. His app shows 2,230 and he is eating close to maintenance, and the scale will confirm it while the numbers on screen insist otherwise.

This is why the macro calculator output should be read as a target range rather than a specification. The difference between two reasonable macro splits is far smaller than the error in measuring either, and Sacks and colleagues demonstrated the practical consequence in the New England Journal of Medicine (360:859-873, 2009): more than eight hundred adults randomised across four diets differing substantially in macronutrient composition showed similar weight loss at two years, with adherence rather than composition predicting the outcome.

What tracking is genuinely good for

Given all of that, the case for tracking rests on three things it does well, none of which requires accuracy.

It is a consistent relative instrument. If your bias is stable, then 2,000 logged kilocalories is reliably less than 2,400 logged kilocalories, even if neither figure is true. That is enough to detect and manage a change in intake, which is the actual job.

It calibrates against your own weight trend. Log consistently for two weeks while weighing daily and reading a rolling seven-day average. If your weight holds steady, your logged intake is your personal maintenance figure in your own units, bias included. Subtract from that number rather than from a calculator's, and the bias largely cancels. This is the single most valuable thing tracking does, and it is the reason calorie needs, equations and calibration treats the calculator figure as a hypothesis to be tested.

It teaches food composition. Most people who track for a month retain a durable sense of what 500 kcal looks like across different foods, and where the protein is. That knowledge persists after the logging stops and is arguably the main long-term benefit.

What tracking cannot do is give you an accurate absolute intake. If you need that, it comes from a metabolic ward, not an app.

Using a biased instrument well

A few habits do most of the work.

Weigh dense foods and estimate the rest. Reuse the same database entries. Log everything, including the things you would rather not, since an omitted item is a much larger error than a mis-measured one. Keep the same convention for raw versus cooked. And accept that restaurant weeks are noisy rather than trying to force precision onto them.

Then hold the whole exercise at the right resolution. Compare week-on-week averages, not daily totals. Give any change three to four weeks before judging it. When your logged intake and your weight trend disagree, the weight trend is the measurement and the log is the estimate, which is the opposite of how most people read it. That is also the first thing to check before concluding anything about your metabolism, as why weight loss plateaus happen sets out.

When to stop tracking

Tracking is a tool, and it is not the right tool for everyone.

For some people, logging becomes the point rather than the method: days are graded as successes or failures on whether the numbers were hit, foods become permitted or forbidden, and eating without recording produces genuine anxiety. That is a recognised pattern, and it is not a discipline problem to be solved by tracking more carefully.

Some signs are worth naming: food and numbers taking up more attention than you want, distress at eating something unlogged, restriction or compensatory exercise following a high day, or a history of disordered eating. If any of that is familiar, the right move is to speak to a clinician rather than to refine the method. There are effective ways to manage intake that involve no numbers at all, and choosing one of them is not a failure of rigour.