PIQ Labs
Essay

Guessing looks exactly like knowing

Food-photo apps identify meals well. The numbers they attach are a different problem, and the interface doesn't tell you which is which.

A photograph of a bowl of turkey chili, submitted to a nutrition app, returns the contents within about a second. Say it reports 620 calories and 41 grams of protein. The figures are illustrative; the format is not. They arrive typeset exactly like any other fact on the screen.

The relevant question is where those figures came from. Sometimes a nutrition database. Sometimes the model itself, on the basis that the words "turkey chili" sit near numbers of roughly that size in its training data. The two cases reach you identically: same typography, same absence of qualification. These interfaces commonly supply nothing that would let you tell them apart.

Users rarely raise this. Complaints about such apps are concrete and mechanical, about what the camera could not see: the cooking oil, the portion. Provenance is seldom questioned, which is unsurprising. A number displayed by a computer has historically been a number the computer looked up.

Fluency is not knowledge

A language model generates text by predicting the next token, and runs the same process whether it is reproducing something or constructing it. It exposes no reliable marker separating the two at the application layer; work on detecting fabrication from model internals is active and unsettled, and no production-grade signal was available to us. There is also a formal result suggesting the problem is not merely engineering debt: facts appearing rarely in training set a floor under the error rate for that class of fact.13 Self-reported confidence does not fill the gap: in our own vision evaluation the model returned a confidence of 0.95 on an item it had hallucinated. We stopped treating that score as sufficient on its own. It still gates escalation below a threshold, but structural signals such as an empty item list now trigger escalation independently of it.

What the measurements show

Some scale first. Google Research's Nutrition5k pipeline, using depth-derived volume under laboratory conditions, reported mean calorie error of 41.3 kcal, 16.5% of the test set's mean.5 That is roughly the field's published ceiling, achieved with hardware better than a phone's and on cafeteria plates averaging under six ingredients. Note that it is an error expressed against the dataset mean, not a per-dish percentage error, which makes it more flattering than the figures below rather than less. Every figure below should be read against it.

The most instructive finding is not that these systems fail to see food. A 2024 evaluation in Nutrients screened 800 apps and tested sixteen; of the seven with AI image recognition, food-component identification ranged from 46% to 97%, and the four that also produced automatic energy estimates identified components at 87–97%.1 Among those four, mean energy error on single-component dishes ran from 47% below the reference value (SD 41) to 44% above it (SD 102), a spread of 91 points between two apps that identified food at 87% and 92%. The standard deviations are as large as the means; the finding is the dispersion, not the decimals.

Identification and nutritional estimation are separable problems, and the second is substantially less solved.

A 2025 study analysed 52 photographs of 12 weighed base dishes and tested three models. Energy error was 35.8% for GPT-4o, 35.8% for Claude 3.5 Sonnet and 64.2% for Gemini 1.5 Pro. For protein the same three scored 60.7%, 61.7% and 109.9%.2 Protein, the macronutrient clinical guidance emphasizes most for patients on GLP-1 medications, was estimated worse than energy by every model tested.

Now the qualification that matters, because it changes the argument rather than softening it. The authors benchmark against four external studies of estimated diet records validated against doubly-labeled water in weight-stable athletes, pooling to 26.5% error (95% CI 20.3–32.8), overlapping GPT-4o's 27.3–45.1. On that basis they judge the models comparable to traditional self-report, and they are right. But a person writing down what they ate knows they are estimating. The interface reporting 620 calories does not say so. The problem was never that the machine is worse than a human guess. It is that a human guess announces itself and this one does not.

A 2026 research letter tested repeatability: twenty meal-image pairs, half varying camera angle and half varying plating, each run sixty times. It concludes that "relying on a single prediction from a single image is insufficient to guarantee estimation reliability." It measured precision only, not accuracy, and two of its authors hold equity in a company applying the technique.3 The meal did not change between images. The reported figure did.

Deployed apps have now been tested against weighed food. Investigators at the NIH's intramural program photographed 102 meals of known composition. All four apps tested underestimated energy, by roughly a third against a mean true meal energy of 918 kcal.4 This is a conference abstract from July 2026, not peer reviewed or published, and its reference standard is national food composition tables applied to the planned meals rather than calorimetry. The investigators' stated caution is worth repeating in their own terms: users of photo-based tracking who do not adjust portions or enter amounts should treat the results accordingly.

We are naming no individual product. The failure is a property of the approach, not of any company's diligence, and it applies to ours wherever a value cannot be sourced. A league table would imply otherwise.

Our own model is not exempt

Before describing what we do about any of this: our photo recognition runs on a small, inexpensive model. A 2026 evaluation in Nutrients tested current-generation models across three datasets and found the small ones markedly worse at estimating energy from images (R² between 0.13 and 0.33, against 0.47 to 0.60 for a mid-sized model), and concludes by recommending the larger one for this task.11 That finding applies to the class of model we use. We are not naming the specific version, because it is a configuration value we expect to change as the economics change; what would not change is the architecture.

Two details of that paper matter, and neither rescues us. Its R² comes from an ordinary least-squares fit with an intercept, so it measures how well a model tracks variation across meals rather than whether its numbers are right; it is unchanged by any systematic over- or under-estimation. And one of its three datasets is packaged-product images evaluated with text-reading suppressed, which is a different task from recognising a plate of food. So the range blends two problems, and the metric cannot see the systematic underestimation the weighed-food work reports.

We chose a small model on evidence, and the evidence is ours, with what that implies. In a 2×2 comparison (a small and a larger model, each working blind and each grounded in the user's own description of the meal) across seven photographs from real use, the grounded condition mattered more than model size, and the larger model was sometimes worse when working blind. We do still escalate to a larger model when a blind read comes back unusable. That is an internal evaluation with n=7, unreviewed, with images and criteria chosen by us, and its grounded conditions supply the model with the answer, so it cannot separate visual identification from transcription of what the user said. It justified an engineering decision and supports no general claim.

The published finding measures energy estimation and ours measured identity, so ours cannot corroborate it; they point the same direction, and the published one carries far more weight. Our response was not to buy a larger model but to stop treating any model's estimate as the answer where a sourced value exists, and to label it plainly where none does. That is a claim about the architecture, and it holds whichever model sits underneath, which is the point of building it that way.

The error is unobservable

In many domains a wrong output announces itself. Defective code fails to compile; a wrong address produces a failed journey. The error surfaces and gets corrected.

A calorie estimate does none of that. You will not discover that the chili was 780 rather than 620. Nothing in your day tells you. The error is silent, and because logging repeats daily it accumulates in whatever direction it leans.

So "roughly right on average" is a weaker defense than it sounds, because individual meals are what people act on. That is an argument against trusting a single meal's absolute figure. It is not an argument against logging. Averaged over many meals a biased instrument can still show a trend, though the repeatability finding above means we cannot promise even that the bias is stable. What survives without qualification is simpler: for most people the useful output of a food log is not the day's total but whether a protein source appeared at each eating occasion, a question that survives a large error in the number beside it.

Relevance to GLP-1 therapy

A great deal is asserted confidently in this area and rather less is established.

Total energy intake falls substantially during treatment: reported reductions of 16% to 39% compared with placebo, mostly from short controlled studies rather than sustained free-living intake.6 Protein tends to hold its share of the plate, so protein grams fall in step. In an exploratory 24-week observational study, protein intake averaged 0.90 g per kilogram of body weight per day (43 patients at baseline, 28 completing, and completers beginning at a lower BMI than those who left), above the general-population allowance of 0.8 g/kg, below the higher intakes suggested during weight loss.7 That figure comes from self-reported dietary assessment, the method whose error this essay has already discussed, which is worth holding in mind before treating it as precise.

What is not disputed is that substantial weight loss by any means costs lean tissue: a 2026 analysis put fat-free mass at 29.8% of total weight lost at six months and 24.8% at twelve, measured by bioelectrical impedance.12 Nor is the response disputed: adequate protein and resistance training, recommended by every side. What is contested is whether these medications make the proportion worse than equivalent weight loss by other means. A post-hoc MRI analysis in patients with type 2 diabetes found thigh muscle volume broadly tracking prediction, though on the Z-score the authors prefer, the decline was modestly larger than predicted, alongside an improvement in muscle fat infiltration.8 Other analyses argue the losses clearly exceed expectation. We take no position on that question. It does not change what to do.

The dietary recommendation deserves its own qualification. A 2026 international expert consensus (a process supported by Nestlé Health Science, which sells nutrition products) grades its own 1.2–1.5 g/kg figure as expert opinion, and states that "a significant lack of direct evidence to guide clinical practice" is what made a consensus necessary.9 A joint advisory from four US societies notes that higher targets "have also been proposed," declines to specify whether they should rest on actual, adjusted or ideal body weight, and cautions that prolonged intake at or above 2 g/kg/day should be avoided. That advisory also states that increased protein alone "is likely inadequate to support the preservation of muscle mass in the absence of structured resistance/strength training."10 Read that as protein being necessary but not sufficient, not as protein being dispensable.

The defensible summary is narrow. Intake falls; the composition of what remains matters more than it did; a patient given a protein target by a clinician has a measurement problem in knowing whether it is met. It is the measurement these systems perform least well.

Not medical advice, and not for everyone. Protein targets belong to your prescriber, and the figures above do not apply to every patient. In chronic kidney disease protein is commonly restricted rather than increased, sometimes below 0.8 g/kg, though on dialysis the target is usually higher rather than lower. Reduced kidney function is also common, and often undiagnosed, among people with type 2 diabetes. Cirrhosis and previous bariatric surgery shift the target too. Do not act on a number from an essay.

Detailed food tracking also is not right for everyone: for someone with a history of disordered eating it can worsen symptoms, and whether to track at all is a decision to make with a clinician beforehand. And there is a point where reduced intake stops being the medication working: severe abdominal pain warrants urgent medical attention rather than a log entry, and inability to keep fluids down or persistent vomiting warrant contacting your prescriber the same day.

Why instruction is insufficient

The apparent remedy is to instruct the model. A prompt-level directive of this kind is common and it was ours: do not state nutrition figures you cannot support.

It fails structurally rather than lexically. Compliance would require the model to notice that a figure in the sentence it is composing came from nowhere, which is exactly the discrimination it cannot make. The instruction is followed in most instances, arguably worse than consistent failure, because intermittent compliance discourages checking. An instruction asks the model to comply; a constraint does not depend on compliance.

What a system would have to do instead

Four requirements follow from the evidence above. They are stated generally because they apply to any product in this category, including ours, and a reader evaluating some other app can use them as a checklist.

The model must be structurally unable to state a figure as though it were sourced. That means composing text with placeholders resolved from the stored record, and inspecting the output as it streams to remove bare figures the model produced on its own. It does not mean the model never supplies a number: where a food is in no database, something has to fill the gap, and the honest design is not silence but a number routed through a channel that renders it visibly as an estimate. The guarantee is labelling, not abstinence.

Every figure must carry its origin, visibly. A value read from a label is not the same kind of fact as one inferred from a photograph, and a total must inherit the weakest source among its parts. We meet this on the meal card and in the data export; the day's running totals carry no marker at all, which is a gap we have not closed. Averaging the provenance of a meal produces a more flattering label and a less accurate one.

It is worth being clear that even the strongest tier is not exact. A nutrition label is a declaration, not a measurement, and federal compliance rules permit the food to differ from it: naturally occurring protein may run as low as 80% of the declared value, and calories as high as 120%.14 Both permitted directions happen to work against someone tracking a protein floor on reduced intake. The tiers rank how a figure was obtained. None of them, including the best, makes it exact.

Quantity must be the user's to correct, and the correction must be arithmetic. Portion is the error term these investigators most often point to, and no model currently judges it well. Stating that half was eaten should rescale the stored numbers by that fraction: ordinary arithmetic on existing values, not a fresh estimate. Photographing the plate afterward should compare it against the original image where one exists, and in either case adjust the stored numbers rather than start over.

A correction must be credited only as far as it goes. This is the constraint that prevents a system flattering itself. When a user restates some of a meal's figures, the untouched values must not be promoted along with them. The record should say which numbers the user supplied and which remain the machine's estimate. In the apps we have examined, one edit is treated as a blessing on the whole entry, and that is how unearned confidence reaches the summary a clinician is handed at the next visit.

We implemented these four in our own product, and the point of stating them as requirements rather than features is that they are checkable against any product, including a competitor's and including ours. The first three are visible from the interface. The fourth is visible by correcting one number in a meal and reading what the record then says: whether it claims the whole entry is now exact, or names which figures you supplied and leaves the rest marked as ours.

What this does not address

So we do not claim to be the most accurate nutrition app. That claim is unfalsifiable, it is made universally, and no user can check it. Our narrower claim is that every meal's figures show their origin, on the card and in the export. Since we have just conceded that no independent evaluation of this app exists, the consistent thing is to invite one. We will supply test meals, our own instrumentation and engineering support to any research group willing to run a weighed-food comparison, and we will publish the result whichever way it goes.

A note on sources

Researching this, search returned several sites presenting app-accuracy figures in the register of published research, converging on roughly 1% error for one product. We could locate no journal, DOI or index record for the principal document, and did not cite it, for a reason needing no investigation. About 1% error would be roughly sixteen times better than 16.5%, the best published laboratory result, obtained with better hardware than a phone. The figure is not optimistic; it is implausible.

Separately, a widely repeated statistic we had intended to use turned out, on checking, to have been welded together from two unrelated papers, one about dialysis patients.

Both were handed to us by automated search tools, without qualification. That is this essay's argument one level up: such systems produce well-formed claims efficiently and have no mechanism for marking which are sourced. The remedy is procedural: follow each figure to its origin, and treat any that cannot be followed as decoration. Every figure here taken from the literature carries a reference below. Several changed while we checked them, and two corrections we made on the advice of a confident reviewer turned out to be wrong, which we discovered by reading the papers.

Conclusion

Provenance does not remove the need for trust. It relocates it, from an unqualified figure to a stated source you can weigh, and, where the number depends on a quantity, to a correction you can make yourself. When a figure looks surprising, you should be able to tell whether it was read from a label or inferred from a photograph, and discount it accordingly.

An honest figure is not always a precise one. Often it is a range, sometimes an admission that a value is unknown. It is the only kind that supports reasoning, and for someone eating substantially less than they used to and trying to keep the muscle they have, reasoning about intake is the substance of the task.

References.
1. Li X, Yin A, Choi HY, Chan V, Allman-Farinelli M, Chen J. Evaluating the quality and comparative validity of manual food logging and artificial intelligence-enabled food image recognition in apps for nutrition care. Nutrients. 2024;16(15):2573. doi:10.3390/nu16152573. PMID 39125452.
2. Fridolfsson J, Sjöberg E, Thiwång M, Pettersson S. Performance evaluation of 3 large language models for nutritional content estimation from food images. Curr Dev Nutr. 2025;9(10):107556. doi:10.1016/j.cdnut.2025.107556. PMID 41081011.
3. Wang Z, Lane D, Waki K. Toward robust AI-assisted dietary assessment for diabetes self-management: quantifying and decomposing large language model prediction variability from meal images. JMIR Diabetes. 2026;11:e102715. doi:10.2196/102715. PMID 42696510.
4. Charles O, Flacke EU, Turner S, Yang S, Airaghi K, Vallone N, Herra L, Darcey VL, Chung ST, Hengist A. Photograph-based AI features in calorie-tracking apps underestimate energy content of meals. Curr Dev Nutr. 2026;10(Suppl 1):108738. doi:10.1016/j.cdnut.2026.108738. Presented at NUTRITION 2026, American Society for Nutrition; conference abstract, not peer reviewed.
5. Thames Q, Karpur A, Norris W, Xia F, Panait L, Weyand T, Sim J. Nutrition5k: towards automatic nutritional understanding of generic food. Proc IEEE/CVF CVPR; 2021:8903–11. doi:10.1109/CVPR46437.2021.00879. arXiv:2103.03375.
6. Christensen S, Robinson K, Thomas S, Williams DR. Dietary intake by patients taking GLP-1 and dual GIP/GLP-1 receptor agonists. Obes Pillars. 2024;11:100121. doi:10.1016/j.obpill.2024.100121. PMID 39175746. Three of four authors are Abbott employees.
7. Babazadeh D, Therrien S, Fitch AK, Steinberg FM. CRAVE study. Obes Pillars. 2026;19:100292. doi:10.1016/j.obpill.2026.100292. PMID 42440974.
8. Sattar N, Neeland IJ, Dahlqvist Leinhard O, et al. Tirzepatide and muscle composition changes in people with type 2 diabetes (SURPASS-3 MRI): a post-hoc analysis. Lancet Diabetes Endocrinol. 2025;13(6):482–93. doi:10.1016/S2213-8587(25)00027-0. PMID 40318682.
9. Sievenpiper JL, Ard J, Blüher M, et al. Nutritional and lifestyle supportive care recommendations for management of obesity with GLP-1-based therapies. Obes Pillars. 2026 Mar;17:100228 (online 11 Nov 2025). doi:10.1016/j.obpill.2025.100228. PMID 41502845. Supported by Nestlé Health Science.
10. Mozaffarian D, Agarwal M, Aggarwal M, et al. Nutritional priorities to support GLP-1 therapy for obesity: a joint advisory from ACLM, ASN, OMA and TOS. Am J Clin Nutr. 2025;122(1):344–67. doi:10.1016/j.ajcnut.2025.04.023. PMID 40450457.
11. Nakagawa S, Yamamoto A. Prompt engineering and model selection for LLM-based nutritional estimation from food images: a multi-dataset investigation. Nutrients. 2026;18(12):2017. doi:10.3390/nu18122017. PMID 42356404.
12. Wang Z, Wang L, Zhang X, et al. Body composition changes after bariatric surgery or treatment with GLP-1 receptor agonists. JAMA Netw Open. 2026;9(1):e2553323. doi:10.1001/jamanetworkopen.2025.53323. PMID 41511769.
13. Kalai AT, Nachum O, Vempala SS, Zhang E. Evaluating large language models for accuracy incentivizes hallucinations. Nature. 2026;653(8116):1047–51.
14. US Food and Drug Administration. Nutrition labeling of food: compliance procedures. 21 CFR 101.9(g).

Nutrician is not a medical device and does not diagnose, treat, cure or prevent any disease. PIQ Labs LLC.