We build AI for decisions where a confident guess is expensive.
A language model produces an answer it looked up and an answer it invented by exactly the same process, and presents them identically. In most software that's a nuisance. In health and in money it is the whole product. So we don't ask our systems to be careful. We build them so that inventing a number isn't something they can do.
AI food apps identify your meal with 87–97% accuracy and then report calorie numbers that differ by 90 percentage points between apps. Seeing the food and knowing its nutrition are separate problems, and only one of them has been solved. What the measurements show, why telling the model not to guess doesn't work, what we built instead — and the parts we haven't fixed.
Read the essay →Figures are resolved from real data and filled in after the model has written its sentence. A gate inspects the output as it streams and removes anything the model tried to state on its own. When the data is missing, the number doesn't appear — the failure mode is silence, not invention.
A value read off a label is not the same kind of fact as one inferred from a photograph, and the interface never lets them look alike. Provenance travels with the number, and a total inherits the weakest source it was built from — averaging would flatter us.
An estimate is labelled as an estimate and shows its range. When the honest answer is that we don't know, the product says so instead of picking something plausible. Plausible is the problem.
We don't claim to be the most accurate. That claim is unfalsifiable and everyone makes it. We claim something narrower and checkable: you can always see where a number came from.