A Hybrid Machine Learning and Deep Learning Framework for Real-Time Air Quality Prediction in Indian Cities: Multi-City Empirical Validation, Ablation Analysis, and Governance Assessment
Keywords:
Attention mechanism, Hybrid machine learning, Deep learning, LSTM, Air quality forecasting, India, CPCB, Reproducibility, Ablation study, Statistical significanceAbstract
Air pollution remains a critical public health and policy challenge in India, where several metropolitan regions rank among the world's most polluted. This paper presents a hybrid Machine Learning / Deep Learning (ML/DL) framework for short-horizon PM2.5 forecasting that combines gradient-boosted tree models (XGBoost, LightGBM) with an attention-augmented Long Short-Term Memory (Attention-LSTM) network via a validation-loss-weighted ensemble. Unlike prior single-city or simulated evaluations, the framework is empirically validated end-to-end on real, publicly available Central Pollution Control Board (CPCB)-derived daily air-quality records (2018–2024) for four Indian metropolitan cities Delhi, Hyderabad, Mumbai, and Chennai comprising 8,257 station-days in total. All results reported here, including RMSE, MAE, MAPE, R², an ablation study, and Diebold–Mariano statistical-significance tests, are computed from actual model training and testing runs rather than illustrative figures. The hybrid ensemble achieves the lowest or near-lowest 1-day-ahead RMSE in three of four cities and reduces error relative to the weakest baseline by 5–15%, while offering materially better robustness at longer (7-day) horizons than the standalone deep model. Permutation-importance analysis confirms that recent PM2.5 history, rolling statistics, and seasonal encodings dominate predictive attribution, consistent with atmospheric persistence and diurnal traffic-emission physics. A condensed regulatory-compliance discussion (GDPR, EU AI Act, NIST AI RMF, CPCB/NAAQS) is retained for completeness, while the bulk of the paper is devoted to the experimental protocol, dataset documentation, and reproducibility artifacts. Complete code, trained-model configurations, and dataset provenance are released to enable independent replication.





