Mynd Healthcare

Home / Scientific Methods

External and Prospective Evaluation

External evaluation tests transfer beyond development data, while prospective evaluation plans observation forward in time.

#Evaluation should be separate from model building

TRIPOD+AI defines evaluation data as data used to estimate a prediction model's performance. It states that these data should be distinct from data used for training, hyperparameter tuning or model selection, with no participant overlap. Its expanded checklist also asks authors to explain data partitioning and internal validation.

Read what each dataset was used for. The word validation can refer to tuning in machine learning, so a label alone may not establish an independent evaluation. This page concerns prediction models and diagnostic studies, not a universal evaluation sequence for every product.

#External evaluation examines an existing model separately

The TRIPOD+AI expanded checklist describes external validation as evaluating an existing model in a separate dataset. It asks how individual predictions were obtained and whether the model was updated during evaluation. It also asks authors to describe data sources, setting, eligibility and representativeness.

Read what was held fixed, which data were separate and whether the model changed. A separate dataset is evidence about the population and conditions it contains, not a permanent approval for all future settings. An external label should not replace a description of the actual comparison.

#Prospective describes the timing of the plan

STARD asks whether data collection was planned before the index test and reference standard were performed, or afterwards. Its explanation notes that prospective and retrospective are used inconsistently. It therefore asks for the actual timing and methods, rather than relying on those words.

Read when the question and collection plan were set. Prospective timing and external evaluation describe different aspects of a study. A prospective label alone does not establish a representative sample, an appropriate reference standard or a benefit to patients.

#Performance needs context and limits

FDA's diagnostic-test guidance explains that appropriate participant selection alone is not sufficient for external validity; device version, users and operating conditions also matter. TRIPOD+AI distinguishes model discrimination, calibration and clinical utility, and asks for performance measures with reasons for their selection.

Read the version, setting, users and outcome alongside the estimate. Evidence of predictive performance is not automatically evidence that using the model improves care. Mynd has performed no external evaluation, prospective study or clinical review on this page.

Source note

The sections above were checked against the linked sources. No clinical review has been performed. This is general research education, not a clinical guideline.