Model Performance Measures
Different performance measures describe different strengths and errors, so no single score can establish whether a clinical model is useful.
#Start with the prediction task
A model predicting a continuous measurement is not evaluated in exactly the same way as a model predicting whether an event occurs. For a future event, the time point matters too. The BMJ external-validation tutorial separates overall fit, calibration and discrimination because they describe different aspects of performance.
Read what outcome was predicted, over what time period and in which population. A headline score without those details leaves the evaluation question unclear. Results from one dataset do not establish performance in every setting.
Evidence: BMJ: evaluating clinical prediction models in external data
#Ranking is not probability agreement
For a binary outcome, discrimination concerns whether people with the outcome tend to receive higher predicted probabilities than people without it. The area under the receiver operating characteristic curve, or AUROC, describes that ordering. It does not tell the whole story about the probabilities themselves.
Calibration asks whether predicted and observed outcomes agree. A model can rank people usefully while estimating risks too high or too low. Read calibration evidence alongside discrimination, with uncertainty estimates. The BMJ tutorial also notes that discrimination depends on the population being evaluated: a score is not a fixed property independent of case mix.
Evidence: BMJ: evaluating clinical prediction models in external data
#Overall fit and decisions answer different questions
Overall prediction error is another part of evaluation. For continuous outcomes, mean squared error compares predicted and observed values. For binary outcomes, the Brier score compares predicted probabilities with observed outcomes. These summaries complement calibration and discrimination rather than replacing them.
Decision analysis asks a further question: how do benefits and harms compare when a prediction leads to an action? Net benefit weighs those consequences using a specified trade-off; decision curves compare strategies across relevant probability thresholds. A high AUROC alone cannot establish a good clinical trade-off. Read which action, alternatives and assumptions were included.
Evidence: BMJ: evaluating clinical prediction models in external data / BMJ: net benefit and decision curves
#A performance report is not deployment permission
IMDRF calls for testing under clinically relevant conditions, including intended populations, subgroups and human-AI interactions. A pooled score cannot describe every group or every effect of a model in a care process. Clear reporting helps readers see what was examined and what remains uncertain.
TRIPOD+AI is reporting guidance for prediction-model studies, not a quality-appraisal tool or permission to use a model. These sources provide reading principles, not a universal score required for approval. No model evaluation service, clinical threshold or validated Mynd device is offered through this page.
Evidence: IMDRF: good machine learning practice principles (2025) / TRIPOD+AI: reporting prediction-model studies (2024)
Source note
The sections above were checked against the linked sources. No clinical review has been performed. This is general research education, not a clinical guideline.