Comparing Machine Learning Algorithms for Flood Prediction
DOI:
https://doi.org/10.70687/869y5616Keywords:
Flood, Prediction, K-Means, Random ForestAbstract
Flooding is a major environmental hazard that requires accurate and reliable prediction to support disaster mitigation and early warning systems. This study applied a combination of K-Means clustering and Random Forest classification to predict flood conditions based on water-level data from DKI Jakarta. The dataset was obtained from Open Data Jakarta and contained observations recorded from January to December 2020. We utilized K-Means to group water-height observations and evaluated cluster configurations from k = 3 to k = 20. The resulting cluster information was then incorporated as an additional feature into the Random Forest model. We evaluated the model using Accuracy, Precision, Recall, F1 Score, and a confusion matrix. The results showed that the K-Means configuration affected the classification performance. The highest F1 Score reached 0.90 at k = 14, while the highest Accuracy reached 0.96 at k = 15 and k = 20. The Random Forest model achieved an Accuracy of 0.95, with weighted Precision, Recall, and F1 Score of 0.96, 0.95, and 0.95, respectively. Class 0 achieved the highest F1 Score of 0.98, while Class 2 obtained a Recall of 0.99. These findings demonstrate that the K-Means and Random Forest approach can effectively support flood prediction using water-level data and provide a promising basis for developing data-driven flood monitoring and prediction systems.
Downloads
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Hamzah, Selly Dwipuspita Ayuningsih (Author)

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.





