Calibration and Decision Thresholds
Calibration concerns the reliability of predicted probabilities, while decision thresholds determine when an output leads to a particular action.
#Reliable probabilities differ from good ranking
Calibration concerns agreement between predicted risks and observed outcome frequencies. As a hypothetical illustration, among a group assigned a risk of 10%, roughly 10 in 100 would experience the specified event if those predictions were well calibrated. This does not tell us which individual will have the event.
Discrimination concerns how well predictions distinguish people with and without the outcome. A model may rank people well while systematically overestimating or underestimating their risks. The calibration methods paper and TRIPOD+AI treat these as separate aspects of evaluation, not interchangeable quality labels.
Evidence: Van Calster and colleagues: calibration of predictive models / TRIPOD+AI: reporting prediction-model studies (2024)
#Inspect more than the average
Comparing the average predicted risk with the overall event frequency checks calibration in the large. A calibration curve examines agreement across the range of predicted risks. A good overall average can still conceal inaccurate estimates in parts of that range.
Calibration curves need enough data to be precise; small samples can make the apparent shape uncertain. TRIPOD+AI recommends reporting performance with uncertainty and in relevant groups. Read the outcome definition, follow-up period, population and plot together before describing probabilities as reliable. No fixed sample-size rule is established by this page.
Evidence: Van Calster and colleagues: calibration of predictive models / TRIPOD+AI: reporting prediction-model studies (2024)
#A threshold connects a prediction to an action
A threshold probability describes the point at which an action is considered worthwhile relative to not acting. The decision-curve guide links it to the relative harms of missing disease and unnecessary intervention. The relevant action might be further investigation or monitoring, not only treatment.
A decision curve compares the net benefit of strategies across a clinically reasonable range of thresholds. It does not choose the correct threshold by locating the model's highest curve. Preferences and consequences must be considered before that comparison. Decision curves are research tools, not direct instructions for an individual patient.
Evidence: Vickers and colleagues: interpreting decision curve analysis
#Recheck when a model or setting changes
Calibration can change when patient populations, event rates, measurements or care practices differ from the development setting. Recalibration or wider model updating may improve estimates, but the change needs to be described and its resulting performance assessed.
TRIPOD+AI asks authors to report model updating and subsequent performance. IMDRF addresses ongoing monitoring and controls for retraining risks. A revised probability or threshold can change who is flagged, so a good historical score is not permission to change a clinical workflow without further evaluation. No clinical threshold is recommended here.
Evidence: Van Calster and colleagues: calibration of predictive models / TRIPOD+AI: reporting prediction-model studies (2024) / IMDRF: good machine learning practice principles (2025)
Source note
The sections above were checked against the linked sources. No clinical review has been performed. This is general research education, not a clinical guideline.