We build AI for decisions where a confident guess is expensive.
A language model produces a figure it retrieved and a figure it constructed by the same process, and presents them identically. In most software this is a nuisance. In health and in finance it is the substance of the product. We therefore do not ask these systems to exercise care; we constrain them so that constructing a figure is not an available behaviour.
Automated food-recognition systems identify meals well — up to 97% of components — then report calorie figures differing by 90 percentage points between applications. Identification and nutritional estimation are separable problems, and only the first is largely solved. A review of the measurement literature, why prompt-level instruction is insufficient, the constraints we implemented, and the limitations that remain — including a published finding unfavourable to the model we deploy.
Read the essay →Figures are resolved from real data and filled in after the model has written its sentence. A gate inspects the output as it streams and removes anything the model tried to state on its own. Where the data is absent, no figure appears; the failure mode is omission rather than fabrication.
A value read off a label is not the same kind of fact as one inferred from a photograph, and the interface never lets them look alike. Provenance travels with the value, and a total inherits the weakest source among its constituents; averaging would yield a more favourable classification and a less accurate one.
An estimate is labelled as an estimate and shows its range. Where the accurate answer is that a value is unknown, the application reports this rather than selecting a plausible substitute.
We do not claim to be the most accurate. That claim is unfalsifiable, it is made universally, and a user has no means of evaluating it. The narrower claim, which can be evaluated, is that every figure displays its origin.