AI Is Entering Disease Surveillance — But Can It Tell Signal From Noise? 🤖🦠
Published in Computational Sciences, Biomedical Research, and General & Internal Medicine
Explore the Research
Page Not Found | PLOS Neglected Tropical Diseases
The promise of intelligent surveillance 🤖
Disease surveillance is becoming increasingly computational. Machine-learning models can analyse large datasets, detect complex associations and potentially identify emerging signals that may be difficult to recognise through conventional surveillance alone. In veterinary medicine and public health, this creates particularly interesting possibilities because laboratory results, electronic health records, genomic information, wildlife surveillance, livestock data and environmental observations could increasingly contribute to interconnected surveillance systems. Recent research already demonstrates growing interest in machine learning for antimicrobial resistance surveillance, early-warning systems and decision support across One Health settings (Rabbani et al., 2026). In veterinary epidemiology, researchers are also exploring AI-based systems that move beyond prediction toward closed-loop surveillance and decision support for animal epidemics (Wang et al., 2026). Yet as these systems become more sophisticated, I find myself returning to a simpler question: what exactly are these models learning from?
A model can learn the test, not only the disease 🧪
In epidemiology, an observed positive result is not always equivalent to true infection. Diagnostic tests differ in sensitivity and specificity, sampling strategies vary considerably, some populations are monitored more intensively than others, and surveillance systems differ between regions, species and institutions. These differences already complicate conventional estimates of disease prevalence, and machine learning does not automatically make them disappear. A sufficiently powerful algorithm may become very good at detecting patterns created by the surveillance system itself. If one population is primarily evaluated using a highly sensitive assay while another is assessed using a different screening method, an algorithm may identify a difference that partly reflects diagnostic methodology rather than biological variation. Similarly, when animals are sampled predominantly because they appear clinically abnormal, a model may learn patterns associated with the selection process rather than population-level disease risk. The resulting model can therefore be statistically impressive while answering a subtly different question from the one researchers originally intended to investigate.
What our own data taught me 🔍
This issue became particularly clear to me while working on our recent study of Brucella exposure in wild canids. Instead of treating every positive serological result as equivalent evidence of infection, we used a misclassification-aware multi-assay framework that accounted for imperfect diagnostic sensitivity and specificity. Across 48 wild-canid serology populations involving 3,925 animals, the estimated global misclassification-adjusted true seroprevalence was 8.2%, whereas confirmed active infection based on PCR or culture was estimated at 3.9% (Sarvestani et al., 2026). The important lesson for me was not simply that one estimate was higher than another. It was that the interpretation of the data changed when diagnostic uncertainty became part of the analysis rather than being treated as background noise. Machine-learning approaches were also used to explore variation associated with factors such as assay type, host species and geographic region, but those analyses became meaningful only when the methodological characteristics of the underlying evidence were explicitly considered (Sarvestani et al., 2026).
More data do not automatically mean better evidence 📊
Artificial intelligence is frequently associated with increasingly large datasets, which can create the impression that scale itself resolves uncertainty. Epidemiology suggests otherwise. More observations cannot necessarily compensate for systematically biased observations. A very large sample collected from an unrepresentative population may still provide a distorted picture of the population of interest, while thousands of diagnostic results produced using tests with different performance characteristics do not automatically become comparable simply because they are combined in the same database. The same problem appears at the level of global surveillance. A recent scoping review examining minimum essential datasets for disease surveillance through a One Health lens found that 89.3% of the included studies primarily focused on human-health information, while animal and environmental interfaces were substantially less represented (Li et al., 2026). This illustrates an important limitation for AI-enabled One Health surveillance: our technical capacity to integrate information may be developing faster than our ability to generate balanced, standardised and genuinely cross-sectoral datasets.
AI cannot make missing information exist 🌍
One Health surveillance is inherently heterogeneous because hospitals, veterinary clinics, farms, wildlife programmes, diagnostic laboratories, genomic platforms and environmental monitoring systems were not originally designed as components of one unified dataset. They measure different variables, use different definitions and collect information at different frequencies and levels of resolution. Li et al. (2026) identified substantial limitations in interdisciplinary and cross-sectoral integration in existing surveillance dataset development, highlighting the importance of standardisation before heterogeneous information can be meaningfully combined. Artificial intelligence may be able to process fragmented information at unprecedented scale, but it cannot automatically resolve missing geographic coverage, inconsistent definitions, poor sampling or absent animal and environmental data. In this sense, one of the greatest limitations of AI-based surveillance may not be the sophistication of the algorithm, but the structure of the evidence supplied to it.
From prediction to useful intelligence 🎯
A model that predicts accurately within one dataset is not automatically a useful surveillance system. For veterinary or public-health decision-making, it is also necessary to understand whether predictions remain reliable in different populations, whether models are appropriately calibrated, whether their inputs are measured consistently and whether their outputs can realistically inform action. This distinction is becoming increasingly visible in current research. In a recent scoping review of machine learning for antimicrobial resistance surveillance, Rabbani et al. (2026) distinguished technical prediction from actionable surveillance intelligence and found important gaps in areas including external validation, temporal stability, prospective evaluation and implementation readiness. This changes the question that should be asked of an AI model. Instead of asking only, “How accurate is the algorithm?”, perhaps we should also ask, “What evidence produced that accuracy, and will the same relationship exist when the model leaves the dataset in which it was developed?”
Veterinary epidemiology may be central to solving this problem 🐾
The growing role of AI in animal disease surveillance makes these questions particularly relevant to veterinary epidemiology. Recent work has proposed increasingly integrated AI systems for epidemic prediction, automated diagnostics and decision support in livestock settings, reflecting how rapidly this field is progressing (Wang et al., 2026). Yet veterinary epidemiologists have long dealt with many of the challenges that AI systems now face: imperfect diagnostics, heterogeneous host populations, biased sampling, incomplete surveillance coverage and uncertainty about whether observed infection reflects exposure, active disease or simply the characteristics of the detection method. These are not peripheral methodological details. They determine the biological meaning of the data. As AI becomes more deeply embedded in animal-health surveillance, epidemiological reasoning may therefore become more important rather than less important.
The next challenge for disease surveillance 🧭
Artificial intelligence will almost certainly become a larger component of infectious-disease surveillance, and its capacity to integrate complex datasets, recognise patterns and support earlier decision-making offers significant opportunities. However, the strongest surveillance systems may not necessarily be those using the most complex algorithms. They may instead be those that combine computational power with traditional epidemiological discipline: representative sampling, validated diagnostics, transparent definitions, external validation, careful interpretation and explicit acknowledgement of uncertainty. Recent One Health research similarly suggests that moving from prediction toward useful surveillance intelligence requires more than model performance alone; it requires harmonised data, validation and clear links between analytical outputs and real-world decisions (Rabbani et al., 2026).
My previous work with prevalence studies and systematic reviews has repeatedly brought me back to the same principle: numbers rarely speak independently of the methods that produced them. Systematic reviews can reveal how diagnostic methods, sampling strategies and geographic gaps shape an evidence base. AI introduces another layer to this problem because algorithms can learn those same structures with remarkable efficiency. The future challenge, therefore, may not simply be whether artificial intelligence can identify more patterns in disease data. It may be whether we can distinguish biological signal from methodological noise before asking machines to learn from both.
References
Li, T., Jia, L., Qiang, N., Zheng, J., Ran, J., Zhang, X. and Han, L. (2026) ‘Progress and challenges in development of minimum essential dataset for disease surveillance through a One Health lens: a scoping review’, Infectious Diseases of Poverty, 15, article 45. doi: 10.1186/s40249-026-01437-6.
Rabbani, S.A., El-Tanani, M., Matalka, I.I., Sharma, S., Saini, M. and Kumar, R. (2026) ‘From surveillance to intelligence: a scoping review of machine learning for antimicrobial resistance surveillance intelligence across One Health’, Frontiers in Public Health, 14, 1922265. doi: 10.3389/fpubh.2026.1922265.
Sarvestani, N., Shams, F., Mirshahi, A., Pato, M., Farbod, A.J., Khayatderafshi, A., Payami, M. and Abdous, A. (2026) ‘From tests to truth: A misclassification-aware machine learning framework for estimating brucellosis seroprevalence in wild canids’, PLOS Neglected Tropical Diseases, 20(3), e0014029. doi: 10.1371/journal.pntd.0014029.10.1371/journal.pntd.0014029
Wang, Y., Zhou, Y., Li, L., Xu, H., Ni, W. and Li, X. (2026) ‘Closed-loop artificial intelligence agents for animal epidemic prediction and decision support in livestock farming: a review’, Frontiers in Veterinary Science, 13, 1968632. doi: 10.3389/fvets.2026.1968632.
AI use disclosure: Generative AI was used to assist with language refinement and structural editing. The scientific content, interpretation and references were reviewed by the author, who takes responsibility for the final published version.