Deep Learning-Based Building Detection from HR Aerial Imagery: Application of Mask R-CNN in Baghdad
DOI:
https://doi.org/10.24996/ijs.2026.67.9.29Keywords:
Building Detection, Computer Vision, Deep Learning, HR Aerial Images, Mask RCNN, SegmentationAbstract
This paper suggests a practical application method that effectively detects buildings in HR aerial photos in Baghdad by creating a segmentation mask using the Mask R-CNN deep learning model. The paper used aerial 10 cm resolution images of Baghdad as the main case study, focusing on a particular urban region neighborhood with various land feature types, building sizes, and population density. Another public access aerial image with a resolution of 7.5 cm was used in New Zealand to assess the DL model's performance for different building architectures. The study found that Mask R-CNN has effectively extracted buildings from HR aerial images of Baghdad, outputting indices with reasonable accuracy. The method achieved a segmentation accuracy of 81.655% and a bounding box accuracy of 83.585% on test data from various areas within the city. Analysis revealed that most incorrect predictions occurred within Baghdad's Al-Sadr district. This was attributed to irregular building layouts, closely spaced or adjacent buildings, and high image brightness caused by sunlight. This issue could be mitigated by considering image acquisition time to minimize the effects of sun angle and its resulting high brightness. Effective processing parameters and image enhancement techniques mitigated occlusion issues and reduced false positives, enabling these methods to facilitate the successful detection of buildings in Al-Mansour and Al-Karrada districts that were not initially present in the ground truth data.




