Sampling and Representativeness
Recruitment and selection determine which people a study describes and how far its findings may extend.
#The intended population is the starting point
FDA's diagnostic-test guidance calls for studying subjects from the intended-use population. STARD asks authors to report eligibility criteria, how potential participants were identified, the setting and the dates of recruitment. These details help explain who the findings concern.
Read the intended population alongside the people actually included. Eligibility alone does not explain who reached the study or who was left out. The sources here concern diagnostic accuracy and prediction models, not one sampling recipe for every health-science study.
Evidence: FDA (March 2007): reporting diagnostic-test evaluation results / STARD 2015: explanation of diagnostic-accuracy reporting items
#Recruitment methods affect who is represented
STARD distinguishes consecutive, random and convenience series. Its explanation notes that enrolling people only on certain days or when an investigator is available can limit representativeness. It also describes how different ways of finding eligible participants can change the range of conditions in the study.
Read how participants entered the study, rather than assuming that a large database is a representative sample. Consecutive recruitment still operates within a particular location and eligibility rule. A sample can be useful for a focused question without representing every population.
Evidence: STARD 2015: explanation of diagnostic-accuracy reporting items
#Missing difficult cases can distort performance
FDA describes spectrum bias when important patient subgroups are missing. Its example is a study containing very healthy people and people with severe disease, while omitting intermediate cases that are harder to diagnose. The guidance also states that increasing the number of subjects does not by itself reduce bias.
Read which disease states, other medical conditions and demographic groups were included. A larger sample can reduce sampling uncertainty while leaving a selection problem intact. This page does not set demographic quotas or a universal minimum sample size.
Evidence: FDA (March 2007): reporting diagnostic-test evaluation results
#Development and evaluation need their own descriptions
TRIPOD+AI asks prediction-model studies to describe development and evaluation data separately, including their sources, suitability and representativeness for the target population. It asks for settings, locations, eligibility and participant dates, rather than a dataset label alone.
Read whether the people, measurements and setting match the claimed use. The same model can face different populations in different settings. Mynd has not recruited participants, assessed a dataset's representativeness or certified a model on this page.
Evidence: TRIPOD+AI (2024): expanded prediction-model reporting checklist
Source note
The sections above were checked against the linked sources. No clinical review has been performed. This is general research education, not a clinical guideline.