OPTIMIZATION OF CARDIOVASCULAR DIAGNOSIS THROUGH ARTIFICIAL INTELLIGENCE APPLIED TO THORACIC RADIOLOGICAL IMAGES: A SYSTEMATIC REVIEW OF DIAGNOSTIC PERFORMANCE, GENERALISABILITY AND TRANSLATIONAL READINESS
DOI:
https://doi.org/10.4238/gwgn0353Keywords:
artificial intelligence; deep learning; chest radiography; computed tomography; cardiovascular diseases; coronary artery calcium; opportunistic screening; external validation; algorithmic fairness; systematic review.Abstract
Background. Cardiovascular disease remains the leading cause of death worldwide, yet risk stratification depends on inputs that are frequently missing in routine care. Chest radiographs and chest computed tomography (CT) are acquired in vast numbers for non-cardiac indications and contain cardiovascular information that is not systematically extracted. Artificial intelligence (AI) has been proposed as a means of recovering that information. Objective. To systematically appraise the diagnostic and prognostic performance, the generalisability, and the translational readiness of AI models that infer cardiovascular status from thoracic radiological images. Methods. Following PRISMA 2020, eight pre-specified queries were executed on 4 September 2026 across three federated evidence-discovery platforms (Consensus, Elicit, scite), which jointly index MEDLINE/PubMed, Scopus, Semantic Scholar and arXiv. Eligible records were primary studies (2019–2026) developing or validating an AI model whose index test was a chest radiograph or chest CT and whose target condition was cardiovascular. Risk of bias was appraised with PROBAST+AI and QUADAS-2 signalling domains. Because of substantial clinical and methodological heterogeneity, synthesis was structured and narrative; a pre-specified exploratory analysis compared paired internal and external discrimination. Studies were mapped onto a two-axis Translational Readiness Matrix combining an adapted imaging-efficacy hierarchy with a five-level generalisability ladder. Results. Of 130 records retrieved and 116 screened, 44 studies were included (chest radiography n = 28; chest CT n = 16). Discrimination clustered by task. Automated coronary artery calcium (CAC) quantification on non gated chest CT was the most mature application, with intraclass correlations against gated reference standards of 0.90–0.99 and risk-category agreement (κ) of 0.65–0.95. Radiograph-based inference was consistently more modest: area under the receiver operating characteristic curve (AUC) 0.73–0.90 for any CAC, 0.79–0.92 for reduced left ventricular ejection fraction and left ventricular structural abnormality, and 0.83–0.92 for moderate to-severe valvular lesions. In six studies reporting paired estimates, discrimination fell on external validation by a median of 0.045 AUC points (range −0.130 to +0.033). Only 11 of 44 studies (25%) reported performance disaggregated by sex, age or ethnicity. Thirty-seven studies (84%) reached only technical or diagnostic-accuracy efficacy; a single study demonstrated therapeutic impact, and none demonstrated patient-outcome or economic benefit. Conclusions. AI can convert routine thoracic imaging into an opportunistic cardiovascular risk signal, and for CAC quantification on chest CT the technical case is essentially settled. The field's binding constraint is no longer accuracy but evidence architecture: external, demographically disaggregated and prospectively collected validation, and outcome-anchored endpoints. We propose the Translational Readiness Matrix as a reporting and commissioning instrument to make that gap explicit.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.

