From Orbit to Leaf: A Critical Systematic Review of Artificial Intelligence for Multi-Scale Crop Monitoring and Soybean Leaf Disease Diagnosis
Keywords:
Precision agriculture, systematic review, critical appraisal, remote sensing, satellite image time series, unmanned aerial vehicles, soybean leaf disease, deep learning, dataset leakage, generalisation, reproducibility, edge deployment.Abstract
Artificial intelligence is now applied to crop monitoring at three very different observation scales: orbital multispectral time series that cover entire districts, unmanned aerial imagery that resolves individual plots, and proximal leaf photographs that resolve individual lesions. These three literatures cite each other rarely, benchmark against different data, and report accuracy in ways that cannot be compared, so the field has accumulated a large number of high scores and very little transferable evidence. This paper presents a critical systematic review that treats the three scales as one evidence base. Following PRISMA 2020, 4,286 records were screened and 84 primary studies published between 2019 and 2026 were retained, of which 41 are orbital, 17 are aerial and 26 are proximal, with soybean leaf disease recognition used throughout as the proximal test case because it is the most densely populated single-crop task in the corpus. Every included study was appraised with AGRI-CRITIC, an eight-dimension rubric introduced here, and summarised by a composite Field Readiness Index (FRI). The synthesis is deliberately unflattering. Median reported accuracy across the corpus is 93.9 percent, but the median FRI is 37.5 out of 100, only 22.6 percent of studies exceed the field-readiness threshold of 60, and reported accuracy is negatively associated with evaluation-set size at a rate of 1.9 percentage points per decade of samples, which is the signature of small-sample optimism rather than of genuine progress. Statistical dispersion is reported by 4 percent of studies, external-site validation by 8 percent, and among the 19 studies that do evaluate outside their training distribution the mean accuracy loss is 17.8 percentage points. Seven recurring failure modes are identified, including a provenance trap in which merged public leaf corpora let a model separate source collections instead of pathologies. The review closes with FIELD-READY, a ten-item reporting protocol with two hard gates, and a three-horizon research agenda for evidence that survives contact with a real field.





