External and Prospective Evaluation
External evaluation tests transfer beyond development data, while prospective evaluation plans observation forward in time.
#Evaluation should be separate from model building
TRIPOD+AI defines evaluation data as data used to estimate a prediction model's performance. It states that these data should be distinct from data used for training, hyperparameter tuning or model selection, with no participant overlap. Its expanded checklist also asks authors to explain data partitioning and internal validation.
Read what each dataset was used for. The word validation can refer to tuning in machine learning, so a label alone may not establish an independent evaluation. This page concerns prediction models and diagnostic studies, not a universal evaluation sequence for every product.
Evidence: TRIPOD+AI: reporting prediction-model studies (2024)
#External evaluation examines an existing model separately
The TRIPOD+AI expanded checklist describes external validation as evaluating an existing model in a separate dataset. It asks how individual predictions were obtained and whether the model was updated during evaluation. It also asks authors to describe data sources, setting, eligibility and representativeness.
Read what was held fixed, which data were separate and whether the model changed. A separate dataset is evidence about the population and conditions it contains, not a permanent approval for all future settings. An external label should not replace a description of the actual comparison.
Evidence: TRIPOD+AI (2024): expanded prediction-model reporting checklist
#Prospective describes the timing of the plan
STARD asks whether data collection was planned before the index test and reference standard were performed, or afterwards. Its explanation notes that prospective and retrospective are used inconsistently. It therefore asks for the actual timing and methods, rather than relying on those words.
Read when the question and collection plan were set. Prospective timing and external evaluation describe different aspects of a study. A prospective label alone does not establish a representative sample, an appropriate reference standard or a benefit to patients.
Evidence: STARD 2015: explanation of diagnostic-accuracy reporting items / STARD 2015: diagnostic-accuracy reporting statement
#Performance needs context and limits
FDA's diagnostic-test guidance explains that appropriate participant selection alone is not sufficient for external validity; device version, users and operating conditions also matter. TRIPOD+AI distinguishes model discrimination, calibration and clinical utility, and asks for performance measures with reasons for their selection.
Read the version, setting, users and outcome alongside the estimate. Evidence of predictive performance is not automatically evidence that using the model improves care. Mynd has performed no external evaluation, prospective study or clinical review on this page.
Evidence: FDA (March 2007): reporting diagnostic-test evaluation results / TRIPOD+AI (2024): expanded prediction-model reporting checklist / TRIPOD+AI: reporting prediction-model studies (2024)
Source note
The sections above were checked against the linked sources. No clinical review has been performed. This is general research education, not a clinical guideline.