Labels and Ground Truth
Reference labels provide targets for training and evaluation, but their meaning and reliability depend on how they are created.
#A label is defined for a task
The ITU-T annotation report describes marking or categorizing data for machine learning. It distinguishes output tasks such as classification, detection and segmentation: the same input can need a different annotation depending on what a model is meant to produce.
Read exactly what the label means and which task it supports. A category assigned to an image and an outline drawn around a region are different targets. The phrase ground truth should not hide the process used to create the reference answer.
#Understand how the reference was made
The report discusses independent annotation, handling disagreement, expert review and annotator training. It presents an informative framework with configurable choices, not a single mandatory staffing pattern for every health-AI project.
Read who labeled the data, what instructions they used and how uncertain cases were handled. Agreement depends on the task and the process. A report should make those choices visible instead of treating the final saved label as self-explanatory.
#Reference standards need to fit the intended use
IMDRF says reference standards should be clinically relevant, well characterized and fit for purpose, with their limitations understood. The rationale for choosing a standard should relate to the device's intended use and environment.
A model's evaluation compares its outputs with that selected reference. Read whether the reference actually answers the intended question and what uncertainties remain. Strong agreement with a chosen label does not by itself establish patient benefit or eliminate limitations in the reference.
Evidence: IMDRF: good machine learning practice principles (2025)
#Keep annotation context attached
The ITU-T report recommends metadata for the annotation process and discusses how consistency criteria vary across tasks. Weng likewise emphasizes provenance and context when reusing clinical data. Knowing how a target was created helps later readers interpret training and evaluation results.
Read whether annotation methods and limitations travel with the dataset. These sources do not provide a universal truth test or authorize a device. This page is research education, not a claim that Mynd supplies clinical labels, validates diagnostic systems or runs an expert annotation service.
Evidence: ITU-T: data annotation specification (2023) / Weng: clinical data quality across its life cycle
Source note
The sections above were checked against the linked sources. No clinical review has been performed. This is general research education, not a clinical guideline.