In this PhD thesis, we address the segmentation of indoor structured environments into planar patches, which we use for localization and map construction. Planar patches represent extracted machine vision information and are well-defined structural elements. We present effective 3D perception methods based on comprehensive point cloud segmentation approaches and specifically designed for processing large indoor spaces. The proposed algorithms are robustly designed as they perform well even for unstructured data and in the presence of irregularities such as occlusions, noise, outliers, or discontinuous data transitions.
The developed methods are adapted to work with various data sets, including combined 2D–3D modalities such as true RGB-D data or point clouds. We use this data for segmentation and to describe three-dimensional space with regular meshes based on tessellation of n-dimensional Euclidean space. In addition, we consider point estimates of features (e.g., normals, curvature, dimensionality features, etc.) that enable filtering and semantic interpretation of the data. To detect coherent 3D regions and eliminate outliers, we introduced connected component analysis based on the investigation of the properties of volume elements.
A new method for point cloud segmentation called EPCC (Evolving Principal Component Clustering) is proposed, which exploits the locality of data in the ordered data stream of a single view. The approach is based on recursive estimation of the parameters of linear prototypes used to describe different clusters of planar and connected surfaces. To improve the results when processing homogeneous surfaces, we introduced two-stage filtering for detecting outliers and took into account a noise model that allows for the compensation of typical uncertainties in depth sensor measurements. We compared the developed algorithm with established segmentation methods, and the proposed approach achieves better results at longer distances where the signal-to-noise ratio is low – even without prior data filtering. In the database used, more than 90 \% accuracy was achieved in detecting flat surfaces, confirming the high efficiency of the method in non-iterative processing of large point clouds.
The dissertation introduces an innovative method for unsupervised segmentation of multiview point clouds, along with an approach for semantic interpretation of structural elements (SE) in interior spaces, such as walls, ceilings, and floors. These approaches rely on density-based object clustering, incorporating advanced color similarity metrics and low-level features to enhance segmentation by removing outliers and improving the detection of sharp geometric structures. The proposed method offers a refined and versatile solution that effectively handles scene complexity and makes a significant contribution to applications in spatial understanding, SLAM (Simultaneous Localization and Mapping), and object recognition.
Finally, the dissertation presents an innovative approach to localization in structured environments, with a focus on visual odometry that leverages planar patch registration instead of traditional methods. The approach is based on high-level features and involves establishing correspondences through 3D–3D and 2D–2D matching using a newly defined similarity metric. A key advantage of the method is its ability to detect and account for arbitrary orientations of planar patches, which reduces degenerative cases and enhances the accuracy of trajectory estimation. Qualitative comparisons demonstrate high efficiency in data registration and improved localization accuracy. The research further shows that a sufficient number of correspondences enables the construction of geometric maps suitable for the navigation of autonomous mobile systems and drones. Additionally, the concept of comprehensive indoor environment reconstruction is introduced, supporting applications in virtual and augmented reality as well as safe path planning.
|