Abstract

Diabetic retinopathy (DR) is the leading cause of blindness among working-aged adults, affecting approximately one in three people with diabetes. It is largely asymptomatic in the early stages of disease, and therefore routine screening is recommended for all people with diabetes. Machine learning (ML) technologies have been validated for the accurate detection of DR and have the potential to increase access and equity to eye screening globally given the scale of diabetes and therefore DR. We conducted a focused review of the published literature regarding ML approaches to the diagnosis and prognosis of DR. We searched Google Scholar and PubMed for open-access articles reporting findings from systematic reviews, diagnostic accuracy studies, real-world studies, health economic modelling studies and qualitative implementation research of ML-assisted DR screening. ML was found to have comparable diagnostic accuracy to ophthalmologists in internal and external validation studies. However, diagnostic accuracy declined in real-world settings due to differences in image quality or diversity when compared with the refined, training dataset. Economic modelling suggests ML is a cost-effective screening tool for DR, particularly in high-income settings; however, the reliability of models is limited by the quality of input parameters. Despite its diagnostic accuracy and cost-effectiveness, adoption and scale-up of ML-assisted DR screening depends on human and organisational factors such as interoperability with existing clinical infrastructure, human-in-the-loop approaches to screening and patient acceptance of ML-assisted care pathways.

Introduction

Diabetic retinopathy (DR) is the leading cause of blindness among working-aged adults globally (Burton et al., 2021). It is the most common microvascular complication of diabetes, affecting approximately one in three people with the metabolic condition (Burton et al., 2021). Recent modelling commissioned by the World Health Organisation reported a diabetes prevalence of over 800 million people worldwide (Zhou et al., 2024). DR has several stages of severity defined using validated scoring systems. One such scoring system is the International Clinical Diabetic Retinopathy Severity Score defining DR as mild, moderate or severe non-proliferative DR, proliferative DR or diabetic macular oedema (DMO) (Cleland, 2023). 

DR is largely asymptomatic in the early stages, therefore routine screening is recommended every 1-2 years for people with diabetes (Leigh et al., 2026). The scale of screening may not be serviceable with existing human graders such as optometrists and ophthalmologists. Indeed, many countries are experiencing workforce shortages amongst the aforementioned eye care providers. It is therefore important to consider digital health innovations which may augment existing screening services to expand access and provision of DR screening globally (Burton et al., 2021).

Machine learning (ML), specifically convolutional neural networks (CNNs), are highly capable of detecting patterns in images corresponding to a medical diagnosis (Sarvamangala et al., 2022). Several ML algorithms have been developed and validated for the detection of DR (Leigh et al., 2025). However, there is emerging evidence to suggest ML can also prognosticate the risk of developing DR or progressing to advanced stages such as proliferative DR or DMO. The aims of this review are to synthesise, critically appraise and establish research priorities for the application of ML tools to prognosticate DR.

Methods

We conducted a focused review of the open-access literature regarding the development, validation, implementation and cost-effectiveness of ML for DR screening, including prognostication. We used PubMed and Google Scholar to access articles included in this review.

Results

Systematic Reviews of ML-Risk Prediction for Diabetic Retinopathy Progression

In recent years, ML algorithms have become more and more powerful, becoming an increasingly important tool in predicting DR by identifying patients who are most at risk of having severe complications. Huang et al. (2025) examined 15 studies on ML algorithms and found that these models have a moderate to excellent accuracy in identifying DR early. An Area Under the Curve (AUC) of 0.70-0.96 indicates excellent identification of patients who will and will not experience a worsening DR. With such high levels of accuracy, it may be possible to achieve early interventions for patients diagnosed with DR, which can successfully delay or alter the progression of the condition (Huang et al., 2025). Using these research models can help healthcare workers to correctly and reliably assess a patient’s conditions and find out the appropriate steps necessary in order to minimise the patient’s harm. Implementing these ML algorithms in clinics could help in reducing vision loss among people with diabetes and in allowing for early diagnosis of DR. Advances in prognostication of DR will also enable personalised medicine, resulting in unique screening intervals designed to fit a specific patient’s condition, allowing for a less invasive treatment that will suit them better, while also reducing the workload on ophthalmologists (Chustecki, 2024). 

However, despite there being many benefits to implementing AI in DR screening, there are still many limitations to consider. Firstly, many existing ML algorithms are extremely prone to bias: in the review by Huang et al. (2025) of 15 different ML algorithms, all 15 were judged to have a high risk of bias. This bias can come from a multitude of factors, such as “inappropriate selection of subjects, inappropriate treatment of continuous variables, inappropriate methods of screening”, as well as containing missing values (Huang et al., 2025). Not only that, but 11 out of the 15 models were internally validated rather than externally validated, weakening their validity. One of the largest issues remains the accountability and responsibility of ML models. Although they are highly accurate, they are bound to make incorrect judgements at some point due to not being infallible. Despite this, they cannot be held responsible for that wrong decision, and it is difficult to place the blame for a fatal error (Chustecki, 2024). Moving forward, further research on AI should incorporate the perspectives of all stakeholders involved, from clinicians to patients, in order to create the best experience possible (Chustecki, 2024). Additionally, using larger and more diverse datasets when training these models will help ensure that there is minimal bias present (Huang et al., 2025). Although AI has become a promising asset to the medical field, it still needs to be refined in order for it to be fully incorporated in the medical setting.

Another systematic review by Usman et al. (2023) involved ML models to predict the progression of DR. They searched IEEE Xplore, PubMed, SpringerLink, Google Scholar and ScienceDirect for studies published between January 2017 and April 2023. Reviews followed the PRISMA guidelines and also used the CHARMS checklist to assess prediction modelling studies. From 2,026 records, 13 studies were included in the final review. The review discussed two applications, Retmarker and RetinaRisk, that could estimate an individual patient’s risk of DR. Retmarker compared retinal photographs taken at different times to detect changes in microaneurysms. RetinaRisk used information such as diabetes duration, gender, HbA1c, blood pressure and current DR severity to estimate the patient’s risk of developing retinopathy. The studies also used traditional methods like logistic regression, random forests and support vector machines. In addition, deep-learning models were used to analyse retinal photographs. The review reported that 71% of models used deep learning. The highest reported performance came from the Tri-SDN model, which combined fundus photographs with information from electronic medical records and achieved 90.6% accuracy, 96.5% sensitivity and an AUROC of 88.8%. The models used several types of predictors, including retinal images, socio-demographic information, clinical information and genetic data. Models that combined retinal images with clinical information would produce better predictions. Common clinical predictors were age, diabetes duration, blood pressure, HbA1c and cholesterol levels (Usman et al., 2023). External validation was the main limitation identified by Usman et al. (2023); Bora et al. (2021) was the only included study to use a completely independent external dataset. Other problems included small or private datasets and different performance measures (Bora et al., 2021).

Overall, these reviews suggest that ML could help clinicians identify patients at greater risk of DR progression. Models combining retinal photographs and clinical information may be more useful because the two sources provide complementary information. Specifically, retinal images can capture changes within the eye, while clinical records describe wider risk factors. However, these systems should support, rather than replace, clinical judgement. Future studies should compare deep-learning models, clinical risk-factor models and combined models in the same adverse patient populations.

Primary Research of ML-Risk Prediction for Diabetic Retinopathy Progression

ML models differ from conventional software in that they are not programmed with fixed rules; instead, they learn by processing large volumes of labelled examples until they can independently recognise patterns and apply them to new cases. In the context of DR, this capacity for automated pattern recognition addresses a critical limitation of conventional screening: current DR screening relies on manual retinal fundus image analysis, in which trained graders identify lesions, including micro-aneurysms, haemorrhages and exudates, to classify disease severity (Qureshi et al., 2021). This process is resource-intensive, susceptible to inter-grader variability and increasingly strained by rising global diabetes prevalence, which is estimated to exceed 800 million cases (Zhou et al., 2024). Artificial intelligence (AI), and DL in particular, has emerged as a promising alternative. Recent studies demonstrate that automated systems can match or exceed the diagnostic accuracy of human graders while substantially reducing processing time, and that longitudinal retinal image analysis may enable not only diagnosis but prognostication of future DR progression (Wang et al., 2024).

ML and DL systems have advanced from solely detecting DR to predicting its progression in the years ahead (Huang et al., 2022). This evolution could help doctors to provide more personalised screening and focus on patients at greater risk. Traditional yearly checkups may not accurately reflect the risk of DR, allowing unnecessary examinations for some patients and delayed detection of progression in others (Dai et al., 2024). Since DR progresses gradually, annual examinations may not detect changes early enough to prevent vision loss. AI models could suggest potential for longer screening intervals and help identify patients who require earlier intervention, which might assist patients in avoiding blindness (Dai et al., 2024).

DeepDR, a DL model, was invented by Dai et al. (2021). It showed an advance in DR-automated screening, because it could be used on a regular computer, and made it easier for patients to be screened (Dai et al., 2021). The system was trained using approximately 400,000 images of the eye and was tested with external datasets to rate its performance (Dai et al., 2021). The machine classified the severity of DR, detected other eye abnormalities and identified the location of its findings. However, it was unable to predict the condition ahead of time, only giving feedback on how to manage one’s risk factors and prevent it from getting worse (Dai et al., 2021). In efforts to improve the machine, Dai et al. (2024) invented DeepDRPlus, which is able to prognosticate when DR would develop or worsen. The model was able to estimate DR progression for up to five years and extended time between screenings from an annual basis to every 32 months (Dai et al., 2024). The system was tested on a C-index scale, where the closer the subject places to one, the more precise it is. It was scored as a 0.846, implying that it is extremely accurate and could possibly help doctors in the future (Dai et al., 2024). While these findings were promising, Huang et al. (2022) noted that many DL models are trained using homogeneous hospital datasets and populations with similar racial demographics, limiting their applicability. This suggests that some may have concerns when accepting AI-assisted DR screening in clinical practice (Huang et al., 2022).

To build this predictive tool, Wang et al. make use of two AI models, each with different functioning mechanisms. The first model, the ResNeXt-50, analyses incoming retinal images by referencing its training dataset to automatically identify structural patterns and correlations (Wang et al., 2024). The second model, the Mask R-CNN, partitions the retinal images into nine distinct sub-regions, enabling the algorithm to precisely isolate and quantify microaneurysms (Wang et al., 2024). In addition, the 12,768 images of data from EyePACS were divided into three major categories: 80% for training, 10% for validation and 10% for testing (Wang et al., 2024). Methodologically, the researchers trained AI models with the primary dataset, then validated and tested them using the remaining images to calculate the sensitivity, specificity and F1-scores (Wang et al., 2024). The empirical results demonstrated that although individual models perform adequately, a mix of the ResNeXt-50 model for resolution 756×756 and the Mask R-CNN model constitutes the optimal system for clinical settings (Wang et al., 2024). Overall, the ResNeXt-50 had a superior performance in terms of diagnosing DR; on the other hand, the Mask R-CNN model was superior at tracking and localising structural alterations over time (Wang et al., 2024). The researchers concluded that the combined ADRPPA framework successfully forecasted a long-term disease progression, yielding a prognostic F1 score of 0.422 (Wang et al., 2024). While an F1 score of 0.422 may appear low compared to standard diagnostic algorithms, it represents an advancement for prognostic modelling. Although not perfect, the study demonstrates how the hybrid algorithm is able to somewhat identify and conduct a prognosis, which has historically challenged unimodal AI systems (Wang et al., 2024). However, several limitations restrict this research. The model relies exclusively on retinal scans, excluding vital systemic variables such as haemoglobin A1c (HbA1c) levels, duration of diabetes and other relevant patient demographics. 

Building upon these foundations, Nderitu et al. (2024) advanced prognostic AI by evaluating a truly multimodal architecture designed to address the exact limitations of no patient data that was observed in the ADRPPA model. Using data from the UK Diabetic Eye Screening Programme (DESP), the study developed deep learning systems (DLSs) to predict referable DR and referable maculopathy over one-, two- and three-year intervals (Nderitu et al. 2024). To achieve this goal, the study compared: a DLS analysing only systemic risk factors such as age and diabetes duration (tabular DLS); and a DLS analysing both retinal scans and clinical data (multimodal DLS) (Nderitu et al. 2024).

Both research papers align in terms of their core conclusion: hybrid approaches yield the highest prognostic accuracy. However, there are contrasts in methodologies. While Wang et al. achieved their optimal results by combining two different models solely on visual data, Nderitu et al. achieved this by combining different types of data. The empirical results of the DESP study showed that the multimodal DLS significantly outperformed the tabular DLS and provided measurable improvements over the image DLS across every time interval. By incorporating the very systemic patient data that Wang et al. lacked, the Nderitu model yielded highly robust AUC metrics, achieving a 0.95 for one-year and 0.85 for three-year progression during internal testing. The combination of these studies yields a definitive key finding regarding the future of AI screening. The research clearly indicates that while unimodal DLS can leverage longitudinal imaging to forecast DR progression to a certain extent, visual data alone is insufficient for optimal accuracy (Wang et al., 2024). Combining historical retinal images with systemic patient risk factors remains an absolute essential prerequisite for maximising reliability, efficiency and viability of prognostic AI models in real-world primary care settings. 

Important findings by Qureshi et al. (2021) further support DL’s high accuracy and efficiency in DR diagnosis. Qureshi et al. assessed the precision and computational cost of a CNN model that classified the severity of DR based on a five-level grading scale (Qureshi et al., 2021). The DL model incorporated an active deep learning (ADL) framework that selected pertinent regions of interest in the sampling data and implemented an image preprocessing procedure to reduce image noise and optimise feature recognition (Qureshi et al., 2021). When tested with 54,000 retinal fundus photographs provided by the EyePACS dataset, the ADL-CNN model demonstrated a mean sensitivity (SE) of 92.20%, a specificity (SP) of 95.10%, an F-measure of 93% and an accuracy (ACC) of 98% (Qureshi et al., 2021). These results meet clinical acceptability and outperform existing DL detection techniques in terms of SP and ACC, though not SE (Qureshi et al., 2021). Additionally, the ADL-CNN performed at efficient rates, taking an average of 25 seconds for data preparation, 13.45 seconds for feature optimisation training, 1.78 seconds for creation of the final out layer, 5.20 seconds for network classifier training and 8.12 seconds for classification (Qureshi et al., 2021). This vastly outpaces manual retinal screening, which takes 12.8 minutes per patient for non-mydriatic imaging and 9.2 minutes per patient for ultra-wide fundus imaging (Qureshi et al., 2021). Nevertheless, important limitations of the model include its reduced performance in interpreting more complicated, low-contrast images of PDR or moderate-to-severe DR (Qureshi et al., 2021).

DL is not only useful for categorising and diagnosing existing DR, but it may also be applicable for prognosticating the probability of developing DR among diabetes patients who do not yet have the disease. Bora et al. (2021) evaluated the performance of DL in DR prognosis by developing two Inception-V3 DL systems (a one-field and a three-field DLS) that predicted the risk of developing DR within two years. The DL systems were subjected to internal and external validation to assess their performance in comparison to or in conjunction with logistic regression models that only analysed risk factors (Bora et al., 2021). Two longitudinal datasets containing colour fundus retinal photographs and patient risk factors were utilised for retrospective analysis: one for model development and internal validation; the other for external validation (Bora et al., 2021). As the external validation dataset only included primary field fundus photographs and glycated haemoglobin as an overlapping risk factor, the three-field DL system and the all-risk factor logistic regression model were unable to undergo external validation (Bora et al., 2021).

Upon analysis, Bora et al. (2021) found that the DL systems exhibited fair prognosis accuracy that exceeded and was independent of risk factor assessment. For DL-only evaluation, the one-field DLS had an AUC of 0.78 in internal validation and 0.70 for external validation, while the three-field DLS demonstrated an AUC of 0.79 in internal validation (Bora et al., 2021). In contrast, risk factor-only analysis using logistic regression yielded a lower accuracy level of 0.72 AUC in internal validation when all available risk factors were considered (Bora et al., 2021). A combination of the DL systems and risk factors, meanwhile, demonstrated substantially higher accuracy compared to risk-factor only assessment but only marginal improvement compared to the DLS in isolation (Bora et al., 2021). For example, the AUC of the combined DLS and all-risk factor evaluation had an AUC of 0.81 for the three-field system and 0.79 for the one-field system, with absolute improvement of about 0.01 to 0.02 for both validation types (Bora et al., 2021). This suggests that the DL systems were capable of independently and effectively predicting DR risk, though their modest performance and the study’s lack of completeness in terms of validation and data imply the need for further investigation and testing (Bora et al., 2021).

Overall, these studies suggest that DL detection methods, although still in their formative stages, demonstrate strong potential in both assessment accuracy and performance efficiency, offering approaches for streamlined screening as well as personalised patient intervention (Bora et al., 2021). DL’s ability to assess patient risk levels as well as the technology’s rapid speed may reduce congestion in clinics and optimise screening for those with higher probability of developing the disease, though the precise cutoff for risk stratification is as of yet unclear (Bora et al., 2021). The DL system could also be integrated with automated texts or telephone alerts to remind high-risk individuals of screening checkups, improving clinical engagement (Bora et al., 2021). Furthermore, risk evaluation and severity classification could allow for targeted treatments or diabetes management, slowing disease progression and preventing onset of blindness-inducing PDR (Bora et al., 2021). 

Manual risk interpretation methods are outdated, with traditional methods relying on general risk calculations which consider age, blood sugar levels and family history, disregarding other statistics that might be the cause but go unrecognised, leading to reduced personalised treatment. Additionally, non-automatic methods take much longer to adapt to new research, meaning that a diagnosis can go much longer without being discovered. ML combats this by producing much more data than conventional methods, including imaging data, patient history and lab results, allowing far faster and more supported conclusions.

DR occurs in around 77-78% of people with type-1 diabetes and around 25-35% of patients with type 2 (Shukla & Tripathy, 2023). According to a meta-analysis by Wu et al. (2021), which reviewed 60 studies containing over 445,000 images and interpretations, ML models demonstrated accuracy between 90-98% when detecting DR from retinal photos. Importantly, this accuracy remained consistent in many other clinical studies (Wu et al., 2021). However, ML is not a flawless system: the quality and accuracy of the system heavily relies on the quality and relevancy of the data fed to it, and there may be some scepticism or hesitation for patients to entrust their health to an AI system, which could lead to less use of the system overall. Overall, ML represents the shift to more personalised, specific treatments of conditions. In the case of DR, this shift could mean identifying the disease years before the symptoms begin to show.

Real-World Risk Prediction for Diabetic Retinopathy Progression Using ML

Despite the advances in screening and treatment, DR remains one of the leading causes of preventable blindness globally (Irodi et al., 2025). With the asymptomatic progression of DR, many patients are unaware that the disease is worsening until their sight has been severely affected (Fung et al., 2022). This disease has motivated researchers to develop AI and DL models that are capable of predicting which patients are at a high risk of disease progression before severe vision loss occurs (Arcadu et al., 2019). In a study about deep convolutional networks (DCNNs) and predicting DR progression, DCNNs were used to analyse each retinal field in seven-field colour fundus photographs (Arcadu et al., 2019). The results of this study indicate that the peripheral retinal fields provide clearer predictive information of the disease worsening than was seen with the central retinal field. These findings are valuable to medical practice for identifying this form of retinopathy because they provide a new way for doctors and clinicians to identify the high risk of DR, allowing for earlier intervention and closer observation before the patient experiences drastic vision loss. 

The screening methods of today are typically used to focus primarily on clear and existing signs of DR, but they may not always be able to predict how quickly a patient’s disease develops and how likely they are to lose their vision (Arcadu et al., 2019). AI models can analyse many data parameters based on photographs of the patients’ eyes, which can be used to predict the likelihood of DR. If these prediction models that leverage ML were to be implemented, this could greatly allow healthcare providers to detect the rate of progression and create more personalised screening schedules, since individuals experience different rates of worsening. 

A recent study about the possible feasibility of AI in medical practice and predicting DR progression shows how a DL system, such as DeepDR Plus, can have drastic effects in prognostication of the disease’s development (Dai et al., 2024). DeepDR Plus was trained with diverse datasets and used over 100,000 fundus images; this data was then studied by the system and was used to aid in creating personalised predictions of asymptomatic progression that went as far as five years (Dai et al., 2024). This suggests that AI could support more personalised screening intervals by identifying which patients can get checked less frequently and which ones are at a higher risk and in need of a closer follow up.  

These two studies show AI can be used to predict development of DR before a patient’s eyesight is severely affected. The Arcadu et al. study showed that DL models can be used to predict the progression of DR by using retinal fundus photographs, while the Dai et al. study improved on those findings by creating a system that is capable of detecting the disease years in advance by using large and diverse datasets. While these studies have provided another reason as to why AI can be useful in medical practice, it would be hard to fully implement it due to various limitations.

Despite recent developments in this technology, multiple sources (Onyeze et al., 2025; Wu et al., 2026) have highlighted that although the technology itself has a very high accuracy – i.e., sensitivity (the model’s ability, as a percentage, to be able to correctly identify cases of DR) and specificity (the model’s ability to correctly exclude false positives from the study as a percentage) – in detecting both non-proliferative and proliferative DR, additional progress is required to integrate this technology into the health systems of low- and middle-income countries, as well as high-income countries. Although only two relevant sources could be found regarding real-world risk prediction across PubMed, both Onyeze et al. (2025) and Wu et al. (2026) conclude that policymakers should both invest in the infrastructure required to integrate this technology and keep equity and lack of bias at the forefront of future studies, emphasising that AI detection of DR must be tested in even more diverse settings to prove its effectiveness as a diagnostic and a prognostic tool.

Both Onyeze et al. and Wu et al. predominantly focused on data collected from studies from the late 2010s and early 2020s. These studies were taken from five different low- and middle-income countries in the case of Onyeze et al., determined using the income group determined by the World Bank “Country and Lending Groups” classification. Wu et al. recorded data across a larger selection of studies in high-income countries, including the USA, Singapore, the UK, the Netherlands, Germany, Italy, India, China, Brazil, Chile, Vietnam, Hungary, Finland and Sri Lanka.

Limitations pointed out by Onyeze et al. include that current AI models are mostly trained on data from high-income countries, and current sample sizes are too small, limiting evidence. Additionally, there is a current lack of reporting on infrastructure (e.g., for automatic referrals or distribution of specialised cameras) and on integration into current health systems in low- and middle-income countries. Onyeze et al. also stated that transparent reporting needs to be enforced and rules developed for subjects whose images are ungradable by the AI. Similarly, Wu et al. conclude that AI detection is not yet close to being proven to be clinically ready.

In conclusion, AI models across multiple studies (Onyeze et al., 2025; Wu et al., 2026) have shown very high sensitivity, with some even reaching as high as 100% sensitivity for vision-threatening DR (VTDR) (Indian study with AIDRSS, Onyeze et al., 2025). In addition, AI models have also shown similarly high levels of specificity, which is something which will remain important as any diagnostic platform, whether human or machine, needs to be able to preclude false positives as effectively as possible to avoid extra costs and unnecessary treatments. However, data is still scarce regarding the effectiveness of AI-based DR prognostication, requiring continued efforts in new studies and data gathering. AI-based prognostication may also be difficult to integrate into current healthcare systems, especially as systematic reviews for prognostication of a large sample size of studies do not currently exist, highlighted by Wu et al. as an area for continued clinical research.

Health Economics of ML-Assisted Diabetic Retinopathy Screening

The global diabetes prevalence has risen, increasing pressure on healthcare services and the demand for DR (Zhou et al., 2024). DR scanning is an effective strategy in the prevention of vision loss, however it requires a significant amount of resources including programme costs, trained graders and specialist referrals (Leigh et al., 2025). ML has risen to be a new and promising approach towards more efficient DR screening practices by the automation of retinal analysis while maintaining high diagnostic quality (Leigh et al., 2025). Evaluating the cost-effectiveness of ML-assisted screening for DR is vital to find out whether these strategies are effective, can improve patient outcomes and efficiently utilise healthcare resources (Leigh et al., 2025).

Economic data analysis from Australia and Singapore show that ML-assisted DR screening is a more cost-effective alternative to current conventional methods. In Australia, a population-based Markov model was used to compare the cost-effectiveness of AI-assisted DR screening to existing screening methods (Hu et al., 2024). The study demonstrated that universally implementing AI-screening would lead to better patient outcomes, earlier detection of DR, reduced workload for ophthalmologists and is a highly cost-effective method compared to modern methods. Additionally, universal implementation of AI-assisted DR screening was projected to save about AU$ 595.8 million in healthcare costs over a period of 40 years, prevent about 38,000 cases of blindness and lead to 172,090 quality-adjusted life years (QALYs) (Hu et al., 2024). Moreover, the study also reported a benefit-cost ratio of 3.96, implying that every dollar invested into AI-assisted screening could lead to almost $4 in health economic benefits (Hu et al., 2024). Similarly, in Singapore, a cost-effectiveness evaluation of the implementation of AI-assisted teleophthalmology-based DR screening was conducted (Xie et al., 2020). The study concluded that an AI-assisted approach to DR screening reduced annual screening costs by roughly 20%, from $77 to $62 spent per patient. This strategy has also led to less work for human graders and maintained diagnostic performance in comparison to conventional screening practices (Xie et al., 2020). These findings suggest that AI-assisted screening is a more cost-effective pathway, leading to lower economic expenses and better allocation of healthcare resources without compromising the quality of DR detection.

Although both studies have demonstrated beneficial overall outcomes, these findings were based on economic modelling rather than real world trials. This means that these analyses are reliant on assumptions regarding AI’s capability, healthcare costs and screening/follow-up uptake (Xie et al., 2020; Hu et al., 2024). Furthermore, both studies took place in high-income countries with high-income healthcare systems. This could limit the overall validity of these studies as conclusions may not be applicable in other countries with less healthcare funding.

AI is as accurate as ophthalmologists for DR screening, yet its cost-effectiveness is crucial for implementation (Leigh et al., 2025). The main factors affecting AI-enabled DR screening’s cost-effectiveness were found to be labour costs for manual DR screening, screening accuracy of AI systems and patient compliance for referrals (AmbleGandhi, 2025). Specifically, this economic benefit is largely driven by reducing inefficient referrals for patients without evidence of DR while improving adherence to follow-up care for those with VTDR (Ahmad et al., 2025).

Researchers have concluded that AI-enabled screening for DR is generally cost-effective across diverse settings; however, these results depend heavily on idealised mathematical models that neglect real-world constraints. Specifically, cost-effectiveness outcomes were directly tied to how many screened patients went on to complete their recommended follow-up care, with programmes achieving stronger referral completion consistently showing better economic returns (Amble & Gandhi, 2025). This underscores that AI’s economic value depends on more than the screening tool itself, particularly when disease is caught early; it also hinges on whether the surrounding healthcare system can successfully get patients through the referral pipeline (Amble & Gandhi, 2025). 

Most studies assumed equal adherence to follow-up between AI and manual grading, treating it as an idealised, equivalent condition rather than modelling real differences; only two of 18 studies modelled actual adherence gaps between AI and manual grading despite evidence that AI screening improves adherence, since patients receive immediate feedback instead of waiting on a remote grader (Leigh et al., 2026). This is corroborated by a pooled meta-analysis of six studies comprising 20,108 patients with diabetes aged 5-67 years – 6,476 assessed using AI and 13,632 assessed by human graders — which found that initial AI screening for DR significantly increased patient uptake of follow-up appointments compared to human grading reported as odds ratios (OR) along with their respective 95% confidence intervals (CI) (OR = 1.89, 95% CI 1.78-2.01, P = 0.00001) (Rahmati et al., 2025). 

Assessing true cost-effectiveness remains difficult due to the multi-faceted challenge of balancing clinical benefit, system efficiency and patient reach. Corporate initiatives highlight extensive commercial deployment, with deep-learning models facilitating over 600,000 screenings in Indian and Thai hospitals by 2024; however, comprehensive peer-reviewed data supporting their economic viability outside of press releases is lacking (Dave, Steinhubl & McQueen, 2026). This gap between commercial-scale deployment and peer-reviewed evidence is compounded by substantial variation in real-world performance across high- and low-income countries, each facing distinct implementation barriers; consequently, DL models for DR screening cannot be evaluated using a universal, one-size-fits-all framework (Dave, Steinhubl & McQueen, 2026). 

Standardised reporting of the use of AI in healthcare is needed as health technology assessment (HTA) and decision-making organisations are increasingly presented with AI-enabled interventions (Elvidge et al., 2024). Participants in the Delphi panel that developed Consolidated Health Economic Evaluation Reporting Standards for Interventions that Use Artificial Intelligence (CHEERS-AI) explained that it improves clarity around AI-related costs by requiring researchers to specify whether financial estimates are exploratory or simply absent from the analysis, given that the predictive algorithms underpinning these interventions are unlikely to be free to use in practice (Elvidge et al., 2024). Therefore, adopting such standardised reporting frameworks is essential to ensure that health economic evaluations reflect true operational costs rather than optimistic commercial projections.  

Sensitivity and specificity are important indicators of how well a diabetic retinopathy screening test performs. Sensitivity measures how effectively a screening method detects patients who have DR, while specificity measures how accurately it identifies those who are free of the disease (Chen, Chan & Chan., 2026). For population-based screening programmes, both measures must be considered to reduce missed diagnoses while limiting unnecessary referrals. Chen, Chan and Chan (2026) also highlight that selecting a screening strategy should not rely solely on diagnostic accuracy, but should also consider referral burden, the consequences of incorrect results and the overall clinical value of the screening programme. Although both measures are essential, there is often a trade-off between maximising sensitivity and specificity (Wang et al., 2024).

The present study examined whether the most diagnostically accurate artificial intelligence model for DR screening was also the most cost-effective. The findings showed that the AI model with the highest diagnostic accuracy was not the optimal strategy when long-term health and economic outcomes were considered (Wang et al., 2024). Instead, an AI model with higher sensitivity (96.3%) and lower specificity (80.4%) was identified as the most cost-effective strategy, achieving the highest health benefits while maintaining an economic value (Wang et al., 2024). These findings suggest that evaluating AI systems using both diagnostic accuracy and cost-effectiveness provides a more comprehensive assessment of their clinical value. The preferred AI performance threshold also varied according to disease prevalence and willingness-to-pay, indicating that healthcare context influences the most economically favourable strategy (Wang et al., 2024).

Wang et al. (2024) conducted a health economic evaluation using real-world data from 251,535 adults with diabetes enrolled in the Lifeline Express DR Screening Programme in China. They assessed 1,100 combinations of AI sensitivity and specificity by integrating a decision tree with a 30-year Markov model to estimate long-term costs and QALYs for different screening strategies (Wang et al., 2024). They measured cost-effectiveness using incremental cost-effectiveness ratios (ICERs) and performed sensitivity analyses to evaluate the robustness of the model. Similarly, Xie et al. (2020) and Lin et al. (2023) applied decision-analytic economic models to compare AI-assisted screening strategies, demonstrating that these methods effectively estimate the long-term clinical and economic impact of diabetic retinopathy screening before implementation in clinical practice.

The study adopted a societal perspective by incorporating direct medical, non-medical and indirect costs, providing a more comprehensive evaluation of the economic impact of AI-assisted diabetic retinopathy screening. Sensitivity and subgroup analyses strengthened the findings by demonstrating consistent results across different scenarios. However, the study was based on a single AI model and healthcare cost estimates from China, limiting generalisability to other healthcare systems. In addition, diabetic macular oedema and other diabetes-related eye diseases were not included in the economic model (Wang et al., 2024). Future research should evaluate the cost-effectiveness of different AI models in diverse healthcare settings, as regional differences in disease prevalence and economic capacity may influence the optimal screening strategy.

Implementation of ML-Assisted Diabetic Retinopathy Screening

Despite DR screening being a valuable tool in early detection of the disease, it is estimated that over half of patients with diabetes mellitus have failed to meet recommended screening guidelines (Bai et al., 2024). Screening guidelines are individualised for patients based on the severity of the condition and diabetes regulation, with intervals ranging from 3-24 months (Krogh et al., 2025). Healthcare settings increasingly face difficulty maintaining both sufficient staffing and time to meet these intervals for the growing diabetic population around the world. The introduction of AI into primary care settings specifically would not only enhance DRS accuracy, but also improve the accessibility and efficiency of DRS for patients (Krogh et al., 2025). Effective implementation of AI models requires understanding and satisfying multiple parties such as ophthalmologists, staff and patients (Bai et al., 2024). When considering the use of AI in screening, it is important to distinguish between allowing AI to lead DRS, letting AI assist medical professionals in DRS and disregarding AI all together in favour of traditional screenings. 

In an effort to gauge the attitudes of patients towards the integration of AI into their screenings, Krogh et al. conducted a questionnaire-based cross sectional study in which they provided 298 patients in Denmark with an AI-assisted DRS and followed the screening with a survey. Based on the responses, the team was able to conclude that the patients were more open to AI-assisted screening (mean score = 3.60/5) in comparison to traditional ophthalmologist-led screening (mean score = 2.78/5), especially when AI was able to increase the efficiency of the visit (Krogh et al., 2025). The findings highlight that although patients may be wary of autonomous AI leading their DRS, human oversight makes the screening more comfortable and trustworthy. Furthermore, AI models are able to reduce the workload placed on ophthalmologists and screening staff by automating image analysis, allowing clinicians to focus on complex cases (Bai et al., 2024). Maintaining human oversight of AI in DRS not only keeps patient trust, but also meaningfully incorporates AI into existing workflows so healthcare staff are not replaced, but instead given assistance.

Patients’ willingness to accept AI depends on a variety of factors including transparency about how AI functions, data use, informed consent and trust in the healthcare system (Chauhan et al., 2025). Patients may be hesitant to accept AI not because they oppose the technology, but instead because they are not familiar with it (Krogh et al., 2025). Educating patients on the role of AI in DRS, and its ability to analyse retinal information under the supervision of medical professionals, may build trust and reinforce that AI serves as a supportive tool in primary care rather than a replacement for healthcare expertise (Krogh et al., 2025). Ophthalmologists and other medical staff should prioritise informed consent and transparency prior to screening in order to maintain integrity and foster better attitudes towards AI models (Chauhan et al., 2025).

Conclusion

There is emerging evidence to suggest ML models have comparable diagnostic accuracy to trained ophthalmologists. Furthermore, advances in prognostic systems show capacity to predict disease progression over multiple years. Our review found real-world performance consistently falls short of controlled validation studies, driven by image quality variability, dataset homogeneity and limited external validation. Health economic analyses suggest ML-assisted screening is broadly cost-effective, particularly in high-income settings, though these conclusions rely on modelling assumptions in the absence of prospective trial data. Implementation research further highlights that technical accuracy alone is insufficient for adoption; clinician trust, workflow integration and patient acceptance are equally important. Future research should prioritise external validation in diverse populations including low- and middle-income countries, prospective real-world effectiveness trials and standardised health economic reporting using frameworks such as CHEERS-AI, before ML-assisted DR screening can be confidently scaled across the globe.

Bibliography

Ahmad, A., Zakiyah, N. and Suwantika, A.A. (2025) ‘Economic evaluations of artificial intelligence implementation in diabetic retinopathy screening’, Pharmacology and Clinical Pharmacy Research, 9(3), p. 273.

Amble, A. and Gandhi, S. (2025) ‘Cost-effectiveness of artificial intelligence-enabled screening for diabetic retinopathy: a systematic review’ [Preprint]. Research Square. Available at: https://doi.org/10.21203/rs.3.rs-7738522/v1

Arcadu, F., Benmansour, F., Maunz, A. et al. (2019) ‘Deep learning algorithm predicts diabetic retinopathy progression in individual patients’, npj Digital Medicine, 2, p. 92.

Bai, P., Cameron, B., Saeedi, A. et al. (2024) ‘Opportunities to apply human-centred design in healthcare with artificial intelligence-based screening for diabetic retinopathy’, Current Opinion in Ophthalmology, 64(4), pp. 5-8.

Bora, A., Balasubramanian, S., Babenko, B. et al. (2021) ‘Predicting the risk of developing diabetic retinopathy using deep learning’, The Lancet Digital Health, 3(1), e10-e19.

Burton, M., Ramke, J., Marques, A. et al. (2021) ‘The Lancet Global Health Commission on Global Eye Health: vision beyond 2020’, The Lancet Global Health, 9(4), e489–e551.

Chauhan, A., Sarkar, D., Verma, G.S., Rastogi, H., Adapa, K. and Duggal, M. (2025) ‘Evaluating trustworthiness in AI-based diabetic retinopathy screening: addressing transparency, consent, and privacy challenges’, BMC Medical Ethics, 26(1), p. 140.

Chen, K.Y., Chan, H.C. and Chan, C.M. (2026) ‘Clinical setting-dependent diagnostic accuracy of artificial intelligence and store-and-forward diabetic retinopathy screening: a systematic review and meta-analysis’, npj Digital Medicine, 9, p. 400.

Chustecki, M. (2024) ‘Benefits and risks of AI in healthcare: narrative review’, Interactive Journal of Medical Research, 13, e53616.

Cleland, C. (2023) ‘Comparing the International Clinical Diabetic Retinopathy (ICDR) severity scale’, Community Eye Health, 36(119), p. 10.

Dai, L., Wu, L., Li, H. et al. (2021) ‘A deep learning system for detecting diabetic retinopathy across the disease spectrum’, Nature Communications, 12, p. 3242.

Dai, L., Sheng, B., Chen, T. et al. (2024) ‘A deep learning system for predicting time to progression of diabetic retinopathy’, Nature Medicine, 30, pp. 584-594.

Dave, D., Steinhubl, S.R. and McQueen, R.B. (2026) ‘Deep learning-enabled diabetic retinopathy screening: a techno-clinical revolution or just more artificial intelligence hype?’, Diabetes Care, 49(3), pp. 381-383.

Elvidge, J., Hawksworth, C. and Avşar, T.S. et al. (2024) ‘Consolidated health economic evaluation reporting standards for interventions that use artificial intelligence (CHEERS-AI)’, Value in Health, 27(9), pp. 1196-1205.

Fung, T.H., Patel, B., Wilmot, E.G. and Amoaku, W.M. (2022) ‘Diabetic retinopathy for the non-ophthalmologist’, Clinical Medicine (London), 22(2), pp. 112–116.

Hu, W., Joseph, S., Li, R. et al. (2024) ‘Population impact and cost-effectiveness of artificial intelligence-based diabetic retinopathy screening in people living with diabetes in Australia: a cost-effectiveness analysis’, eClinicalMedicine, 67, p. 102387.

Huang, X., Wang, H., She, C. et al. (2022) ‘Artificial intelligence promotes the diagnosis and screening of diabetic retinopathy’, Frontiers in Endocrinology, 13, p. 946915.

Huang, H., Wu, Y., Ye, H. et al. (2025) ‘Risk prediction models for diabetic retinopathy: a systematic review’, Frontiers in Endocrinology, 16, p. 1556049.

Irodi, A., Zhu, Z., Grzybowski, A. et al. (2025) ‘The evolution of diabetic retinopathy screening’, Eye, 39, pp. 1040-1046.

Krogh, M., Germund Nielsen, M., Byskov Petersen, G. et al. (2025) ‘Patient acceptance of AI-assisted diabetic retinopathy screening in primary care: findings from a questionnaire-based feasibility study’, Frontiers in Medicine, 12, p. 1610114.

Leigh, J., Drinkwater, J., Turner, A. et al. (2025) ‘Health economic considerations for the implementation of artificial intelligence-enabled diabetic retinopathy screening: a review’, Clinical & Experimental Ophthalmology, 54(1), pp. 144-161.

Leigh, J.A., Sherrington, A., Barber, A.R.J. et al. (2026) ‘Referral uptake after diabetic retinopathy screening with artificial intelligence-assisted care pathways: a systematic review and meta-analysis’, npj Digital Medicine, 9, p. 468.

Lin, S., Liu, H., Chen, Y. et al. (2023) ‘Artificial intelligence in community-based diabetic retinopathy telemedicine screening in urban China: cost-effectiveness and cost-utility analyses with real-world data’, JMIR Public Health and Surveillance, 9, p. e41624.

Nderitu, P., Nunez do Rio, J.M., Webster, L. et al. (2024) ‘Predicting 1, 2 and 3 year emergent referable diabetic retinopathy and maculopathy using deep learning’, Communications Medicine, 4, p. 167.

Onyeze, C.N. et al. (2025) ‘Effectiveness of AI-based tools in detecting diabetic retinopathy in low- and middle-income countries: a systematic review of diagnostic performance and implementation feasibility’, Cureus, 17(11), p. e96554.

Qureshi, I., Ma, J. and Abbas, Q. (2021) ‘Diabetic retinopathy detection and stage classification in eye fundus images using active deep learning’, Multimedia Tools and Applications, 80(8), pp. 11691-11721.

Rahmati, M., Smith, L. and Piyasena, M.P. et al. (2025) ‘Artificial intelligence improves follow-up appointment uptake for diabetic retinal assessment: a systematic review and meta-analysis’, Eye (London), 39(12), pp. 2398-2406.

Sarvamangala, D.R. and Kulkarni, R.V. (2022) ‘Convolutional neural networks in medical image understanding: a survey’, Evolutionary Intelligence, 15(1), pp. 1-22.

Usman, M., Saheed, Y., Nsang, A. et al. (2023) ‘A systematic literature review of machine learning based risk prediction models for diabetic retinopathy progression’, Artificial Intelligence in Medicine, 143, p. 102617.

Wang, V.Y., Lo, M.T., Chen, T.C. et al. (2024) ‘A deep learning-based ADRPPA algorithm for the prediction of diabetic retinopathy progression’, Scientific Reports, 14, p. 1.

Wang, Y., Liu, C., Hu, W. et al. (2024) ‘Economic evaluation for medical artificial intelligence: accuracy vs. cost-effectiveness in a diabetic retinopathy screening case’, npj Digital Medicine, 7, p. 43.

Wu, J.H., Liu, T.Y.A., Hsu, W.T., Ho, J.H. and Lee, C.C. (2021) ‘Performance and limitation of machine learning algorithms for diabetic retinopathy screening: meta-analysis’, Journal of Medical Internet Research, 23(7), p. e23863.

Wu, Z. et al. (2026) ‘Effectiveness of screening modalities for early detection of diabetic retinopathy: a systematic review and meta-analysis of tele-ophthalmology, AI-based tools, and conventional methods’, Frontiers in Medicine, 13, p. 1778534.

Xie, Y., Nguyen, Q.D., Hamzah, H. et al. (2020) ‘Artificial intelligence for teleophthalmology-based diabetic retinopathy screening in a national programme: an economic analysis modelling study’, The Lancet Digital Health, 2(5), pp. e240-e249.

Zhou, B., Rayner, A., Gregg, E. et al. (2024) ‘Worldwide trends in diabetes prevalence and treatment from 1990 to 2022: a pooled analysis of 1108 population-representative studies with 141 million participants’, The Lancet, 404, pp. 2077-2093.