PIQ Labs
Essay

Guessing looks exactly like knowing.

Photograph a bowl of turkey chili and an AI food app will tell you what's in it. Six hundred and twenty calories, forty-one grams of protein. The answer arrives in about a second, phrased with the same easy confidence it would use for the boiling point of water.

The question worth asking is where that number came from.

Sometimes it came from a real nutrition database. Sometimes it came from the shape of the sentence — from the fact that the words "turkey chili" tend to sit near numbers like those in the text the model learned from. On your screen those two answers are identical. Same font, same certainty, no asterisk. Nothing in the interface gives you what you would need to tell them apart.

Most people never think to ask, and I don't blame them. When users complain about these apps they say sensible, concrete things — it can't see the oil, it doesn't know my portion. Almost nobody asks the question this essay is about, because software has never given anyone a reason to. A number on a screen has always just been a number. That is the assumption worth dismantling.

Fluency is not knowledge

A language model produces text by repeatedly predicting what comes next. That is the entire mechanism. It does not have one mode for recalling a fact and another for inventing one — the same process generates both, and generates both equally well. There is no internal marker separating them, which means there is nothing for an app to surface even if it wanted to.

When a person is unsure, it usually leaks. They hedge, they slow down, they say they think so. A model's hedging is also generated text. It is a writing style, not a measurement of its own confidence. A model can express total certainty about a figure it produced from nothing, in well-formed and reassuring prose.

The output of a lookup and the output of a guess are the same kind of object: fluent text. Any difference between them has to be manufactured deliberately, outside the model.

What the measurements show

This is not a theoretical worry, and the most striking finding is not that the models can't see. They can. A 2024 evaluation of seven AI food apps in Nutrients found food recognition accuracy of 87–97%. The apps knew what the food was. Their energy estimates ran from 47% below the reference value to 44% above it — and two apps that tied on recognition landed 91 percentage points apart on the calorie number.1

Seeing the food and knowing its nutrition are separate problems, and the second one is where the invention happens.

It gets more specific. A 2025 study photographed 52 meals, weighed every component on a calibrated scale, and compared three general-purpose models against that ground truth. Mean absolute percentage error for energy was about 36% for GPT-4o and Claude 3.5 Sonnet. For protein it was 60.7% and 61.7%, and 109.9% for Gemini 1.5 Pro.2 Protein — the number a person on a GLP-1 medication is most often told to watch — was estimated worse than calories by every model tested.

Two caveats, because this essay would be absurd without them. Those are mid-2024 models, tested with deliberately minimal prompting and no database grounding. And the authors' own conclusion is more generous than my framing: they judge the better models roughly comparable to traditional self-reported dietary assessment and describe them as showing promise as screening tools. Their stated limit is that these models are "not yet suitable for precise dietary assessment in clinical or athletic populations where accurate quantification is critical." I would add only that the study's own comparator has humans doing better — estimated diet records at 26.5% error against the models' 35.8%.

Then there is reproducibility, which is the finding I find hardest to argue with. A 2026 study in JMIR Diabetes photographed the same physical meals from different viewpoints and ran each image sixty times. Its conclusion: a single prediction from a single image "is insufficient to guarantee estimation reliability."3 The meal did not change. The answer did.

The shipping consumer apps have now been measured too. Researchers at the NIH's intramural program photographed 102 meals prepared in a metabolic kitchen with every ingredient weighed to a tenth of a gram, then ran four popular calorie apps against them. All four underestimated: Appediet by about 252 kcal per meal, MyFitnessPal 327, Lose It! 333, Cal AI 345 — roughly a third low.4 This is a July 2026 conference abstract and has not completed peer review, and the meals came from a high-fat trial, which is these apps' worst case. I include it because it is the only weighed-food test of the actual products, and I would rather show you its status than leave it out.

For scale: Google Research's own benchmark, using depth sensors under laboratory conditions, reported 16.5% mean calorie error.5 That is roughly the ceiling of what the technique can currently achieve with better hardware than your phone.

The error is silent

In a lot of domains, being wrong announces itself. Bad code fails to compile. A wrong address gets you lost. A wrong claim about history gets contradicted in the replies. The error surfaces and you correct it.

A calorie count does none of that. You will almost certainly never discover that the chili was seven hundred and eighty rather than six hundred and twenty. There is no feedback, no correction, no moment where the mistake becomes visible. It is silent, and because you log every day it compounds quietly in whatever direction it happens to lean — and the NIH data suggest it leans consistently in one direction.

Which is why "roughly right on average" is a weaker defence than it sounds. You do not eat the average.

Why this matters more on a GLP-1

Here I want to be careful, because a great many confident things are said about GLP-1 medications and nutrition, and rather less is settled than the internet suggests.

What is well established: on these drugs people eat substantially less — reported reductions in total intake run from about 16% to 39%.6 When total intake falls, protein tends to hold its share of the plate, so protein grams fall roughly in step with everything else. In a 24-week study of people starting semaglutide or tirzepatide, protein intake averaged 0.90 g per kilogram per day: above the basic recommended allowance, below the 1.2–1.5 g/kg that clinical groups increasingly suggest during weight loss.7

What is genuinely contested is the muscle question. Roughly a quarter to a third of the weight lost on these drugs is lean tissue rather than fat — but dieting does much the same thing, with lifestyle-only weight loss measured around 26%. An MRI study of tirzepatide found thigh muscle loss almost exactly in line with what that much weight loss predicts, and muscle quality improved.8 Other 2026 reviews argue the losses exceed physiological expectation. Researchers disagree, in print, right now.

And the recommendation everyone repeats deserves its own footnote. Every major society statement suggests higher protein alongside resistance training during GLP-1 therapy — but the 2026 international expert consensus grades its own 1.2–1.5 g/kg figure as expert opinion, and says plainly that "a significant lack of direct evidence" is what made a consensus necessary.9 As of this writing no randomized trial has tested protein intake in people taking these medications; the first is still recruiting. The joint advisory from four US societies is blunter still: increased protein alone "is likely inadequate to support the preservation of muscle mass in the absence of structured resistance/strength training."10

So the honest version is narrower than the usual one. Not these drugs destroy muscle and protein saves it, but: you are eating much less, the composition of what remains matters more than it used to, your clinician may well have given you a protein number, and you would like to know whether you are hitting it. That last part is a measurement problem — and it is exactly the measurement the models are worst at.

None of the above is medical advice. Protein targets are a conversation with your prescriber; the same advisory that sets a floor also warns against sustained intake at or above 2 g/kg/day.

Why telling the model not to doesn't work

The obvious fix is to instruct it. Every serious team has some version of do not state nutrition figures you cannot support in their prompt. We did too.

It doesn't hold, and the reason is structural rather than a matter of wording. To follow that instruction the model would have to notice that a particular number in the sentence it is composing came from nowhere — and noticing that is precisely the capability it lacks. You are asking it to detect the one thing it cannot detect. The instruction gets followed most of the time, which is arguably worse than never, because it teaches everyone to stop checking.

An instruction is a request. What we wanted was a constraint.

What we built instead

Three rules, enforced in code rather than asked for in a prompt.

One: the model is not allowed to write a number. It writes sentences with holes in them — placeholders where a figure belongs. The holes are filled afterward by the server, from the stored record of what you actually logged. Meanwhile a gate reads the model's output as it streams and strips any figure it tried to author on its own. If the underlying data isn't there, the number doesn't appear: the sentence gets rephrased, or it doesn't get said. The failure mode is silence rather than invention, and silence is a failure we can live with.

Two: every number carries where it came from. Internally each value is tagged with how it was obtained:

LABELRead directly off the packaging.
KNOWNMatched to an entry in a reference nutrition database.
CORRECTEDYou told us the real value, and you outrank us.
REFERENCE_ESTIMATEDerived from a comparable reference food.
MODEL_ESTIMATEOur weakest tier — an informed guess, presented as one.

A meal inherits the worst tier of anything in it. A salad with a labelled dressing and an eyeballed portion of chicken is a guess, because the eyeballed chicken dominates the error. Averaging the tiers would produce a better-looking number and a less honest one.

Three: uncertainty is shown rather than smoothed. An estimate is labelled as an estimate and carries its range. Where the truthful answer is that we don't know, the product says so instead of quietly selecting something plausible. Plausible is the problem.

Notice what this implies about the research above. Those studies measured what happens when a model's guess is the answer. In our design it never is — the model's job is to work out what the food is, which is the part the measurements show these systems are good at, and the numbers come from lookup. When we genuinely cannot source something, the result is tagged as our weakest tier and shown with a range, because a photo-derived estimate deserves exactly that much trust.

What this does not fix

This is usually where the product turns out to have solved everything. It hasn't, and an argument about honesty that overstated its own results would be a strange thing to publish.

So we don't claim to be the most accurate nutrition app. That claim is unfalsifiable, everybody makes it, and you have no way to check it. We claim something narrower that you can check: every number shows where it came from, and our guesses are visibly distinct from our facts.

A footnote from writing this

Researching this essay, I went looking for evidence on how accurate AI calorie apps really are. Search turned up a tidy network of authoritative-looking sites — a research "initiative," a clinical report, a benchmark, a rankings page — citing one another and converging on a meta-analysis with an impressive study count and sample size. It concluded that one particular app achieves accuracy within about one percent.

None of it exists. No journal, no DOI, no index record; the flagship paper contradicts itself on its own sample size and is self-published by a site carrying an affiliate disclosure for the app it crowns. The giveaway was arithmetic rather than diligence: about 1% error is fifteen to thirty times better than Google's own laboratory result with depth-sensing hardware. It is not a good number. It is an impossible one.

Separately, a widely-repeated statistic I nearly used in this essay turned out to have been welded together from two unrelated papers, one of them about dialysis patients.

Both were handed to me, confidently and without hesitation, by AI search tools. Which is the argument of this essay one level up: the machinery is very good at producing well-formed claims and has no mechanism for telling you which ones are sourced. The only defence I know is the boring one — follow every number to its origin, and treat any figure that cannot be followed as decoration.

The point isn't trust

Most software in this category asks you to trust it. We would rather you didn't have to. The whole design exists so that trust is unnecessary — so that when a number looks surprising, you can find out in one tap whether it was read off a label or inferred from a photograph, and decide for yourself what it is worth.

An honest number is not always a precise one. Often it is a range, or an admission. But it is the only kind you can reason with, and when you are eating substantially less than you used to and trying to keep the muscle you have, reasoning with it is the entire job.

References. 1. Li et al. Nutrients 2024;16(15):2573. 2. Fridolfsson et al. Curr Dev Nutr 2025;9(10):107556 (PMID 41081011). 3. Wang, Lane & Waki. JMIR Diabetes 2026;11:e102715. 4. Hengist et al., NIDDK/NIH — NUTRITION 2026 conference abstract, not yet peer reviewed. 5. Thames et al., Nutrition5k, CVPR 2021 (Google Research). 6. Christensen et al. Obes Pillars 2024;11:100121 (PMID 39175746). 7. Babazadeh et al., CRAVE study, Obes Pillars 2026;19:100292. 8. Sattar et al. Lancet Diabetes Endocrinol 2025;13(6):482–493 (SURPASS-3 MRI). 9. Sievenpiper et al. Obes Pillars 2026;17:100228 — expert consensus, supported by Nestlé Health Science. 10. Mozaffarian et al. Am J Clin Nutr 2025;122(1):344–367, joint advisory of ACLM, ASN, OMA and TOS.

We build Nutrician, a nutrition coach for people on GLP-1 medications, where the rules above are the product rather than a policy page. It is free to try, and it will tell you when it doesn't know.