<?xml version="1.0"?>
<rdf:RDF xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/"><rdf:Description rdf:about="https://repozitorij.uni-lj.si/IzpisGradiva.php?id=113182"><dc:title>LEARNING OF OBJECT REPRESENTATIONS THROUGH ROBOTIC MANIPULATION</dc:title><dc:creator>BEVEC,	ROBERT	(Avtor)
	</dc:creator><dc:creator>Ude,	Aleš	(Mentor)
	</dc:creator><dc:subject>active perception</dc:subject><dc:subject>interactive learning</dc:subject><dc:subject>foveated vision</dc:subject><dc:subject>active vision</dc:subject><dc:subject>autnomousobject learning</dc:subject><dc:subject>object recognition.</dc:subject><dc:description>Objects represent basic blocks of the world in which robots and humans exist. Object
perception is therefore a key capability of intelligent robotic systems. In this thesis
we propose a paradigm that realizes the concept of active exploration for learning of
visual representations for object recognition. We adopt the definition, where an object
is defined as a physical entity, which can be manipulated and who’s properties move
according to the motion imposed on the object.
The capability of monitoring and exploring the environment in a wide field of view,
while retaining high acuity in the foveal view, is achieved by utilizing a system that
uses two actuated cameras per eye. To make 3D reconstruction possible, it is necessary
to model the motor system of the cameras and estimate the camera alignment with the
robot’s kinematic chain. We propose a calibration method to estimate the location
of the cameras on the robot. The robot can then generate sparse point clouds from
triangulated features by using constraints of epipolar geometry.
We developed a floor detection and learning method that extracts the floor appearance
model and masks it in the subsequent images. This prevents point clouds where
most features appear on high contrast floor and few features are found on the objects.
We propose to generate a map of the observed scene by joining consecutive point
clouds into a discrete, equally spaced, 3D grid of voxels. Noise accumulation in the
reconstructed map is tackled by using a proposed point cloud filtering method. In order
to align point clouds in a common coordinate system, we presented a method that fuses
3D sensor information with inertial data to facilitate point cloud alignment and robot
positioning.
The robot detects novel object candidates by processing the estimated point cloud
in peripheral views. Object hypotheses are generated by searching for surface regularity
and feature proximity in the form of geometric structures such as planes, spheres
and cylinders. The robot needs additional information to verify or discard the hypotheses
about object existence. The information comes from motion induced by a human
teacher or by the robot itself while interacting with the object using its manipulation
capability, e. g. by applying pushing actions. By assuming that the object is rigid,
the robot looks if the pushed candidate moved as a rigid body. This way the robot
can confirm (or reject) the object hypothesis and determine features that belong to the
object.
The selected object candidate found in the peripheral view is also inspected in
detail in the foveal view. We designed a controller that directs the cameras towards the
hypothesis’ centroid, ensuring that the object candidate is in the center of the foveal
view. Feature proximity is used to generate the object candidate in foveal views. Since
the foveal view covers a small area, the rigid body motion constraint can be relaxed and
all matched features that exhibit motion can be considered confirmed object features.
After the robot learns some information about the object, it also becomes feasible
to grasp it. We devised a method to systematically observe the object from different
views by using successive grasp-rotate-release action cycles to build the object representations.
Tactile information is used to detect the contact with the object and control
the fingers during the grasp. If the grasping action has succeeded, the robot starts rotating
the grasped object to observe it from different viewpoints and accumulate more
information about the object.
Our final goal is to generate an object representation suitable for object recognition.
The robot applies a number of manipulation actions to the selected object candidate to
accumulate a sufficient amount of data. The aim of these manipulation actions is to
move the object so that it becomes visible from all viewpoints. We propose to build
an object representations consisting of individual object snapshot, separately for the
peripheral and foveal images. SIFT descriptors are used to describe confirmed object
points and we apply the bag-of-features model to create object representations out of
the confirmed features. Support vector machines are then used to train a classifier for
object recognition.
Evaluation of the proposed scientific contributions was performed using two different
robotic platforms, a dual-arm humanoid platform called Kukanoid and a drone.</dc:description><dc:date>2019</dc:date><dc:date>2019-12-11 10:20:01</dc:date><dc:type>Doktorsko delo/naloga</dc:type><dc:identifier>113182</dc:identifier><dc:language>sl</dc:language></rdf:Description></rdf:RDF>
