Missing proper calibration in medical prediction models can cause misrepresented and misguided risk estimates for patients, a research letter suggests, thus degrading model performance. Seth Bergstedt, MS, from the Minneapolis Heart Institute, and colleagues, discussed the corresponding data in a manuscript published online Wednesday in JACC: Cardiovascular Interventions. “We were working on revising the PROGRESS-CTO score using a larger data set and more advanced machine-learning methods, when we noticed that on these new data, the common risk scores appear to be poorly calibrated,” Larissa Stanberry, PhD, from the Minneapolis Heart Institute and corresponding author of the paper, told CRTonline. “Calibration is an important property of the prediction model, though it's often glossed over with most attention spent on improving the model discrimination.” The investigators noted that calibration is quantified by the mean calibration (MC), which is the average difference between predicted and observed risks. The MC can reveal the overall bias. At perfect calibration, the MC should be 0. An MC <0 shows that the risks are overestimated and an MC >0 shows that the risks are underestimated. In this study, the team examined the calibration in 4 risk models used for predicting success in the percutaneous coronary intervention (PCI) procedure for chronic total occlusion (CTO). The 4 risk models were the J-CTO (Multicenter CTO Registry in Japan) score, clinical and lesion-related (CL) score, PROGRESS-CTO (Prospective Global Registry for the Study of CTO Intervention) score and EuroCTO (CASTLE) score. External validation data was taken from the PROGRESS-CTO registry (42 centers, 2015-2022) and used according to the Transparent Reporting of a Multivariable Prediction Model for Individual Prognosis or Diagnosis guidelines. The study cohort consisted of 81% men and participants had a median age of 65 years. Upon evaluation, the estimated MC’s were 0.377 for J-CTO (calibration slope=0.31), 0.140 for the CL score (calibration slope=0.49), -0.061 for PROGRESS-CTO (calibration slope=0.39) and 0.022 for CASTLE (calibration slope=0.89). On average, risk was underestimated among this group of prediction models. The authors also noted that the risk classifications were significantly different between the 4 models (bootstrap p<0.001). Overall, this study indicates that prediction models are not necessarily calibrated, and this can lead to degradation model performance. Dr. Stanberry told CRTonline about the outlook for studying prediction models and risk assessments. “The growing availability of rich clinical data and the democratization of technology are empowering researchers to develop clinical prediction models. With this letter, we hope to draw the attention of the research community to the importance of the calibration in model evaluation and the need of following the robust model development and validation framework when assessing the model performance,” she said. “The rate of data accumulation limits the rate of model updates, so proper validation and testing are here to ensure that the developed models perform reliably and for the foreseeable future.” Source: Bergstedt S, Karacsonyi J, Brilakis ES, et al. Lack of calibration degrades the performance of clinical prediction models: A case of CTO scores. JACC Cardiovasc Interv. 2025 September 17 (Article in press). Image Credit: woravut – stock.adobe.com