An interpretable vision transformer framework for joint wheat leaf disease classification and pesticide recommendation

I am pleased to share my open-access research on an interpretable Vision Transformer for joint wheat leaf disease classification and pesticide recommendation, integrating Grad-CAM explainability for transparent precision-agriculture decision support
Like

Share this post

Choose a social network to share with, or copy the URL to share elsewhere

This is a representation of how your post may appear on social media. The actual post will vary between social networks

Explore the Research

Springer International Publishing
Springer International Publishing Springer International Publishing

An interpretable vision transformer framework for joint wheat leaf disease classification and pesticide recommendation - Discover Artificial Intelligence

This study proposes a novel multi-task Vision Transformer (ViT) framework for simultaneous wheat leaf disease classification and pesticide recommendation. The proposed framework incorporates Grad-CAM interpretability to visualize disease-affected regions and enhance model transparency and trustworthiness. RGB wheat leaf images collected from three public Kaggle datasets were used, covering six disease classes and six pesticide categories. Data augmentation techniques expanded the dataset from 10,720 to 11,408 images. The model was evaluated using fivefold cross-validation with accuracy, precision, recall, F1-score, ROC-AUC, and confusion matrices. Experimental results achieved accuracies of 92% ± 0.01 on original images and 94% ± 0.01 on augmented images, demonstrating stable and effective multi-task learning. Comparative benchmarking with ResNet50 and EfficientNet further confirmed the superior performance of the ViT-based framework. Furthermore, Grad-CAM visualizations improved interpretability by highlighting infected leaf regions relevant to model predictions. The proposed framework demonstrates strong potential for precision agriculture applications by providing accurate disease diagnosis, transparent decision support, and pesticide recommendations. Future work will focus on model optimization and deployment on resource-constrained edge devices.

From Wheat Leaf Images to Smarter Disease Management

Wheat diseases can seriously affect crop productivity and create challenges for farmers who need to identify diseases quickly and decide on appropriate treatments. With the growing use of artificial intelligence (AI) in agriculture, image-based disease detection offers a promising way to support faster and more informed decision-making. Our study explored how AI could be used not only to identify wheat leaf diseases, but also to provide pesticide recommendations within the same framework.

The study introduces a multi-task Vision Transformer (ViT) framework designed to perform two related tasks: wheat leaf disease classification and pesticide recommendation. By combining these tasks, the goal was to move beyond simply recognizing a disease and toward providing more actionable information that could support precision agriculture.

How did we conduct the study?

We used 10,720 wheat leaf images from three publicly available datasets. To improve the diversity of the training data and help the model learn from different variations in leaf appearance, we applied image augmentation techniques. After augmentation, the dataset increased to 11,408 images.

The core of our approach was a Vision Transformer (ViT). Unlike traditional convolutional neural networks, Vision Transformers use an attention mechanism to learn relationships between different parts of an image. This makes them particularly interesting for agricultural images, where disease symptoms may appear in different regions of a leaf.

We also incorporated Explainable AI (XAI) into the framework using Grad-CAM. A major challenge with AI-based agricultural systems is that a model may provide a prediction without clearly showing why it made that prediction. Grad-CAM helps address this issue by highlighting the areas of an image that contributed most strongly to the model's decision. In our case, this allows us to visualize disease-relevant regions of wheat leaves and provides greater transparency in the prediction process.

We compared the proposed framework with established deep-learning models, including ResNet50 and EfficientNet, to evaluate its performance and understand the potential benefits of the transformer-based approach.

What did we find?

The results were encouraging. The proposed framework achieved 92% ± 0.01 accuracy using the original images. After applying data augmentation, the accuracy increased to 94% ± 0.01.

These results suggest that expanding the diversity of the training images can help improve the model's ability to recognize wheat leaf diseases. The comparison with ResNet50 and EfficientNet also demonstrated the potential of combining transformer-based learning with multi-task prediction.

However, accuracy was not the only important consideration in this research. For agricultural applications, it is also important to understand how an AI system reaches its predictions. The Grad-CAM visualizations provide an additional layer of information by showing disease-relevant regions of the leaf. This can make the system's predictions easier to interpret and potentially more useful for researchers and agricultural practitioners.

Why does this research matter?

The motivation behind this work is to contribute to the development of AI-assisted precision agriculture. A system that can identify wheat diseases and provide treatment-related recommendations could potentially help users make faster and more informed decisions.

There is still considerable work needed before such systems can be used reliably in real-world agricultural environments. Models need to be evaluated on images collected under different field conditions, including variations in lighting, backgrounds, disease severity, and wheat varieties. Nevertheless, our findings demonstrate the potential of combining Vision Transformers, multi-task learning, and Explainable AI in a single framework.

For us, an important lesson from this research is that developing an effective AI model is not only about achieving high accuracy. It is also about making the predictions understandable and ensuring that the technology can ultimately provide useful information to the people who may rely on it.

One of the interesting aspects of this work was preparing the image data and developing a model that could learn from different visual patterns associated with wheat diseases. Agricultural images can vary considerably because of differences in disease appearance, image quality, lighting, and leaf conditions. This made data preparation and augmentation an important part of the study. We also explored different deep-learning approaches to understand how the proposed Vision Transformer compared with established architectures such as ResNet50 and EfficientNet. Integrating Grad-CAM added another important dimension to the research, because it allowed us to look beyond the final prediction and examine which regions of the wheat leaf influenced the model's decision. This process helped us better understand the model's behaviour and reinforced the importance of combining performance with interpretability when developing AI systems for agriculture.

DOI: 10.1007/s44163-026-01896-8

Follow the Topic

Agriculture
Life Sciences > Biological Sciences > Agriculture
Artificial Intelligence
Mathematics and Computing > Computer Science > Artificial Intelligence

Related Collections

With Collections, you can get published faster and increase your visibility.

Transforming Education through Artificial Intelligence: Opportunities, Challenges, and Future Directions

Artificial Intelligence (AI) is rapidly changing the educational field by enabling personalized learning, intelligent tutoring systems, automated assessments, learning analytics, and administrative automation.

This collection invites original research, systematic reviews, and visionary perspectives on the transformative impact of AI in education. It aims to explore how AI technologies can enhance equity, inclusion, and efficiency in educational settings across different contexts, including higher education, K-12, vocational training, and lifelong learning. This collection will address technical, pedagogical, ethical, and policy aspects, fostering interdisciplinary perspectives and evidence-based insights.

This Collection supports and amplifies research related to SDG 4 and SDG 9.

Keywords: Artificial Intelligence, AI in Education, Educational Technology, Data Analytics, AI Ethics

Publishing Model: Open Access

Deadline: Nov 30, 2026

AI-driven Ensemble Learning and Feature Engineering for Complex Data

As artificial intelligence (AI) continues to expand into diverse real-world domains, the complexity, volume, and variability of data present new challenges for AI model performance, interpretability, and generalization. This collection focuses on the development and application of AI-driven ensemble learning techniques and feature engineering strategies to address these challenges, particularly in high-dimensional, noisy, imbalanced, and multi-source datasets.

We invite contributions that explore novel ensemble architectures, including stacking, boosting, bagging, and hybrid models, as well as advanced feature selection, fusion, and transformation methods. The collection aims to bridge theoretical innovation with practical deployment, showcasing how ensemble learning and feature engineering can enhance AI model accuracy, robustness, and explainability across domains such as cybersecurity, education, healthcare, smart cities, and industrial systems.

Topics of Interest Include (but are not limited to):

- AI-based ensemble learning frameworks for classification, regression, and multi-label tasks

- Feature selection and fusion techniques for high-dimensional or noisy data

- Handling imbalanced datasets using sampling and ensemble strategies

- Optimization-enhanced ensemble models (e.g., PSO, WOA, GA)

- Fuzzy logic and uncertainty modeling within ensemble systems

- Interpretability and explainability in ensemble-based AI models

- Applications in intrusion detection, student performance prediction, biometric estimation, and smart infrastructure

- Comparative studies and benchmarking of ensemble methods on complex datasets

This Collection supports and amplifies research related to SDG 9.

Keywords: AI-driven Ensemble Learning; Multi-Label Learning; Feature Engineering in AI; Feature Selection and Fusion; Imbalanced Data Handling; Optimization-enhanced Ensembles; High-dimensional Data; Multi-source Data Integration; Hybrid AI Models; Fuzzy Logic in AI Systems

Publishing Model: Open Access

Deadline: Jan 31, 2027