What this tool does
A photo is a noisy ingredient list. The model names what it can see, matches USDA rows when the names are ordinary, and totals the plate. Notes exist because cameras do not see the butter in the pan or the drink off-frame. This is not a lab scan of the food.
How it works
Vision proposes items and portions. Those strings go through the same matching path as a typed meal. Volume from a single angle is a guess — a deep bowl of rice looks like a shallow one. Write weights in the notes when you know them.
How to interpret the result
If the photo is dark, cropped, or a pile of mixed sauce, expect more “estimated” rows. A clean plate of chicken, rice, and broccoli with “2 tsp oil” in the notes will beat a glamorous restaurant shot every time. Use the number to see protein, not to prosecute a 17 kcal rounding error.
Example
A lunch photo of a palm-size chicken piece, a fist of rice, and green beans, with the note “1 tbsp teriyaki, no extra oil” might print ~450–550 kcal and 35–45 g protein depending on the cut. Add “fried in 1 tbsp oil” and the same photo should jump ~120 kcal. If it does not, edit the oil line yourself.
Common mistakes
- Shooting from a dramatic angle that hides half the bowl.
- Leaving drinks, bread baskets, and shared appetizers out of the frame and the notes.
- Uploading a menu cover instead of the plated food.
- Treating a 3% confidence match as a weighed serving.
When this estimate may be off
- Mixed stews, curries, and blended soups with no recipe.
- Low light, motion blur, or a filter that changes color.
- Several people’s plates in one image.
- Packaged food with a Nutrition Facts label you should just read.
