Mynd Healthcare

Home / Health Data & AI

Labels and Ground Truth

Reference labels provide targets for training and evaluation, but their meaning and reliability depend on how they are created.

#A label is defined for a task

The ITU-T annotation report describes marking or categorizing data for machine learning. It distinguishes output tasks such as classification, detection and segmentation: the same input can need a different annotation depending on what a model is meant to produce.

Read exactly what the label means and which task it supports. A category assigned to an image and an outline drawn around a region are different targets. The phrase ground truth should not hide the process used to create the reference answer.

#Understand how the reference was made

The report discusses independent annotation, handling disagreement, expert review and annotator training. It presents an informative framework with configurable choices, not a single mandatory staffing pattern for every health-AI project.

Read who labeled the data, what instructions they used and how uncertain cases were handled. Agreement depends on the task and the process. A report should make those choices visible instead of treating the final saved label as self-explanatory.

#Reference standards need to fit the intended use

IMDRF says reference standards should be clinically relevant, well characterized and fit for purpose, with their limitations understood. The rationale for choosing a standard should relate to the device's intended use and environment.

A model's evaluation compares its outputs with that selected reference. Read whether the reference actually answers the intended question and what uncertainties remain. Strong agreement with a chosen label does not by itself establish patient benefit or eliminate limitations in the reference.

#Keep annotation context attached

The ITU-T report recommends metadata for the annotation process and discusses how consistency criteria vary across tasks. Weng likewise emphasizes provenance and context when reusing clinical data. Knowing how a target was created helps later readers interpret training and evaluation results.

Read whether annotation methods and limitations travel with the dataset. These sources do not provide a universal truth test or authorize a device. This page is research education, not a claim that Mynd supplies clinical labels, validates diagnostic systems or runs an expert annotation service.

Source note

The sections above were checked against the linked sources. No clinical review has been performed. This is general research education, not a clinical guideline.