Efficient Prediction of Water Quality Index (WQI) Using Machine Learning Algorithms

Published in Civil Engineering

Like

Share this post

Choose a social network to share with, or copy the URL to share elsewhere

This is a representation of how your post may appear on social media. The actual post will vary between social networks

Every impactful research project has a story, and for our team, this story began with the question: how can technology improve water quality monitoring to ensure better health and environmental outcomes? Our paper, "Efficient Prediction of Water Quality Index (WQI) Using Machine Learning Algorithms," addressed this question and earned the 2022 Best Paper Award from Human-Centric Intelligent Systems.

The Process and Methodology

The foundation of our research was built on a comprehensive analysis of water quality data sourced from India's diverse water bodies. The dataset included essential parameters such as dissolved oxygen (DO), biological oxygen demand (BOD), pH, and total coliform (TC). To ensure a reliable and replicable process, we designed a robust workflow for data preparation and modeling, as depicted in Figure 1.

Working diagram of proposed model.

Figure 1: Research Workflow
This figure illustrates the sequence of steps followed in the study:

  1. Data Collection: Acquiring datasets from Kaggle, focusing on key water quality parameters.
  2. Data Preprocessing: Addressing missing data using Random Forest imputation and applying Min-Max normalization for scaling.
  3. Feature Selection: Identifying critical variables using a correlation matrix.
  4. Machine Learning Models: Training and testing five algorithms (Neural Network, Random Forest, Multinomial Logistic Regression, Support Vector Machine, and Bagged Tree Model).
  5. Performance Evaluation: Comparing model accuracies and identifying the best performer.

This structured approach not only streamlined our study but also ensured replicability, a cornerstone of rigorous research.

Key Findings and Insights

The performance of the machine learning algorithms was assessed using metrics such as accuracy and kappa values. 

  • The Multinomial Logistic Regression (MLR) model achieved the highest accuracy of 99.83%, setting a benchmark for water quality prediction systems.
  • Random Forest (RF) followed closely with an accuracy of 98.99%, demonstrating its strength in handling complex datasets.
  • Other models, including Neural Network (98.65%), Bagged Tree Model (98.99%), and Support Vector Machine (96.98%), also performed well, though slightly lower than MLR.

The chart underscores the reliability of MLR in WQI prediction, making it an ideal choice for real-world applications.

Practical Implications

Our study's results provide a roadmap for developing efficient, data-driven systems for water quality monitoring. The insights gained can support policymakers, environmental agencies, and researchers in implementing proactive measures to ensure safe water access.

Looking forward, we aim to build a software application using our proposed model, enabling real-time water quality predictions. Such a tool could revolutionize water resource management, particularly in regions facing acute water quality challenges.

Final Thoughts

Winning the Best Paper Award has been a tremendous honor, motivating us to continue exploring the potential of machine learning in solving critical environmental problems. We extend our heartfelt thanks to the editorial board of Human-Centric Intelligent Systems for this recognition and to our research team at VRD Research Lab for their dedication and collaboration.

Follow the Topic

Soil and Water Protection
Technology and Engineering > Civil Engineering > Environmental Civil Engineering > Soil and Water Protection

Related Collections

With Collections, you can get published faster and increase your visibility.

Applications and Challenges of Blockchain Technology in User-Centric Intelligent Systems

This special collection aims to explore the transformative potential and inherent challenges of integrating blockchain technology into user-centric intelligent systems, a rapidly evolving interdisciplinary domain at the intersection of decentralized computing, human behavior modeling, and intelligent analytics.

As human-centric systems increasingly rely on vast, sensitive, and behavior-rich data, blockchain offers promising solutions for trust, transparency, privacy preservation, and decentralized governance. However, its integration into intelligent systems that model, predict, and respond to human behavior introduces unique technical, ethical, and usability challenges.

This collection invites original research, reviews, and case studies that address the following themes:

• Blockchain for Trust and Privacy in Human-Centric Systems: Mechanisms for decentralized trust, identity management, and privacy-preserving data sharing in systems that model user behavior and community dynamics.

• Decentralized User Modeling and Personalization: Blockchain-enabled frameworks for secure and transparent personalization, recommendation, and behavioral analytics.

• Smart Contracts and Autonomous Agents: Applications of smart contracts in automating user-centric interactions, decision-making, and system governance.

• Blockchain in Social and Behavioral Computing: Use of distributed ledgers to track, validate, and analyze social influence, community evolution, and behavioral dynamics.

• Security and Ethical Challenges: Addressing disinformation, misinformation, fairness, and explainability in blockchain-powered intelligent systems.

• Integration with Mobile and Social Sensing: Blockchain applications in ubiquitous sensing environments for healthcare, mobility, and societal impact.

• Scalability and Usability Issues: Technical limitations and human factors affecting the adoption of blockchain in intelligent systems.

• Trustworthy and Explainable AI via Blockchain: Use of blockchain to provide auditability, provenance, and verifiable explanations for AI models and decisions in human-centric systems, enhancing user trust and regulatory compliance.

• Federated and Decentralised Learning for User Privacy: Blockchain-supported federated, swarm, or split learning frameworks that enable collaborative AI model training across distributed users and organisations without exposing sensitive personal or behavioural data.

This collection aligns with the journal’s mission to advance human-centric intelligence by fostering multidisciplinary research that bridges blockchain technology, AI, behavioral modeling, and social computing. Contributions should emphasize both theoretical insights and practical implementations that enhance the understanding and development of secure, ethical, and user-aware intelligent systems.

This Collection supports and amplifies research related to SDG 9 (Industry & Innovation).

Publishing Model: Open Access

Deadline: Dec 14, 2026

Federated Learning and Behavioral AI for Human-Centric Systems

This special collection aims to explore the intersection of federated learning (FL) and behavioral artificial intelligence (AI) in the context of human-centric intelligent systems. As digital interactions increasingly generate vast, sensitive, and distributed behavioral data, there is a growing need for privacy-preserving, decentralized, and ethically aligned AI methodologies. Federated learning offers a promising paradigm by enabling collaborative model training across decentralized data sources without compromising user privacy. When combined with behavioral AI, it opens new frontiers in understanding, modeling, and predicting human behavior in a secure and responsible manner.

This collection invites original research, reviews, and case studies that address theoretical foundations, algorithmic innovations, system architectures, and real-world applications of federated learning and behavioral AI in human-centric domains. Contributions should emphasize privacy, personalization, fairness, explainability, and ethical considerations in the design and deployment of decentralized, collaborative, and privacy-preserving intelligent systems that interact with or model human behavior.

Beyond incremental advances, this collection aims to establish Federated Behavioral AI as a unified research direction by articulating its core problem formulations, system constraints, and evaluation principles. We particularly encourage submissions that formalize the unique characteristics of behavioral data, such as temporal evolution, contextual dependency, multimodality, social influence, and non-stationarity within federated environments, and that propose standardized benchmarks, taxonomies, or reproducible evaluation protocols to consolidate the field.

________________________________________

Topics of Interest Include (but are not limited to):

  • Privacy-preserving behavioral modeling and world model learning
    • Federated personalization and user modeling through learned world representations
    • FL in healthcare, education, mobility, and smart cities, with world models enabling
    • predictive and adaptive human-environment interaction
    • Cross-device and cross-silo federated learning architectures for distributed world
    • model construction and refinement
    • World models for anticipatory reasoning in federated settings: enabling agents to
    • simulate, plan, and act across human-centric domains without centralizing sensitive data

  • Foundation Models and Human-Centric Modeling
    • Behavioral trajectory analysis and prediction
    • Cognitive and affective modeling in decentralized settings
    • Social influence and community behavior modeling
    • Behavioral simulation and digital twin systems
    • Federated fine-tuning of LLMs for behavioral reasoning
    • On-device adaptation of foundation models
    • Federated alignment of large-scale behavioral models
    • Knowledge distillation in federated behavioral systems

  • Security, Trust, and Ethics in Federated Behavioral AI
    • Differential privacy and secure multi-party computation
    • Fairness-aware federated algorithms
    • Explainability and transparency in behavioral AI
    • Governance frameworks for ethical AI deployment
    • Federated robustness against behavioral manipulation
    • Adversarial attacks on behavioral models
    • Secure aggregation for behavioral intelligence
  • System Design and Evaluation
    • Scalable and efficient FL frameworks for behavioral data
    • Benchmark datasets and Standardized federated behavioral metrics
    • Real-world deployments and case studies
    • Human-in-the-loop and participatory AI systems
  • Federated Optimization and System Challenges
    • Communication-efficient and adaptive FL
    • Handling non-IID and behavioral drift detection and mitigation
    • Edge intelligence for behavioral data analytics
    • Resource-aware FL for wearable/mobile devices
  • Applications and Use Cases
    • Mental health and well-being monitoring
    • Personalized learning and adaptive education
    • Smart homes and ambient intelligence
    • Crisis response and behavioral risk detection
    • Intelligent Transportation Systems and Decentralized Autonomous Systems

This Collection supports and amplifies research related to SDG 9 (Industry & Innovation)

Publishing Model: Open Access

Deadline: Dec 02, 2026