This study introduces a multi-task Vision Transformer (ViT) framework for joint wheat leaf disease classification and pesticide recommendation, enhanced with Grad-CAM-based Explainable AI to visualize disease-relevant regions and improve prediction transparency. Using 10,720 images from three public datasets, expanded to 11,408 through augmentation, the framework achieved 92% ± 0.01 accuracy on original images and 94% ± 0.01 after augmentation. Comparative evaluation with ResNet50 and EfficientNet demonstrates the potential of integrating transformer-based learning, multi-task prediction, and explainability for accurate and actionable decision support in precision agriculture.
DOI: 10.1007/s44163-026-01896-8