Search CORE

6 research outputs found

Semantic and generic object segmentation for scene analysis using RGB-D Data

Author: Lin Xiao
Publication venue: Universitat Politècnica de Catalunya
Publication date: 01/01/2018
Field of study

In this thesis, we study RGB-D based segmentation problems from different perspectives in terms of the input data. Apart from the basic photometric and geometric information contained in the RGB-D data, also semantic and temporal information are usually considered in an RGB-D based segmentation system. The first part of this thesis focuses on an RGB-D based semantic segmentation problem, where the predefined semantics and annotated training data are available. First, we review how RGB-D data has been exploited in the state of the art to help training classifiers in a semantic segmentation tasks. Inspired by these works, we follow a multi-task learning schema, where semantic segmentation and depth estimation are jointly tackled in a Convolutional Neural Network (CNN). Since semantic segmentation and depth estimation are two highly correlated tasks, approaching them jointly can be mutually beneficial. In this case, depth information along with the segmentation annotation in the training data helps better defining the target of the training process of the classifier, instead of feeding the system blindly with an extra input channel. We design a novel hybrid CNN architecture by investigating the common attributes as well as the distinction for depth estimation and semantic segmentation. The proposed architecture is tested and compared with state of the art approaches in different datasets. Although outstanding results are achieved in semantic segmentation, the limitations in these approaches are also obvious. Semantic segmentation strongly relies on predefined semantics and a large amount of annotated data, which may not be available in more general applications. On the other hand, classical image segmentation tackles the segmentation task in a more general way. But classical approaches hardly obtain object level segmentation due to the lack of higher level knowledge. Thus, in the second part of this thesis, we focus on an RGB-D based generic instance segmentation problem where temporal information is available from the RGB-D video while no semantic information is provided. We present a novel generic segmentation approach for 3D point cloud video (stream data) thoroughly exploiting the explicit geometry and temporal correspondences in RGB-D. The proposed approach is validated and compared with state of the art generic segmentation approaches in different datasets. Finally, in the third part of this thesis, we present a method which combines the advantages in both semantic segmentation and generic segmentation, where we discover object instances using the generic approach and model them by learning from the few discovered examples by applying the approach of semantic segmentation. To do so, we employ the one shot learning technique, which performs knowledge transfer from a generally trained model to a specific instance model. The learned instance models generate robust features in distinguishing different instances, which is fed to the generic segmentation approach to perform improved segmentation. The approach is validated with experiments conducted on a carefully selected dataset.En aquesta tesi, estudiem problemes de segmentació basats en RGB-D des de diferents perspectives pel que fa a les dades d'entrada. A part de la informació fotomètrica i geomètrica bàsica que conté les dades RGB-D, també es considera normalment informació semàntica i temporal en un sistema de segmentació basat en RGB-D. La primera part d'aquesta tesi se centra en un problema de segmentació semàntica basat en RGB-D, on hi ha disponibles les dades semàntiques predefinides i la informació d'entrenament anotada. En primer lloc, revisem com les dades RGB-D s'han explotat en l'estat de l'art per ajudar a entrenar classificadors en tasques de segmentació semàntica. Inspirats en aquests treballs, seguim un esquema d'aprenentatge multidisciplinar, on la segmentació semàntica i l'estimació de profunditat es tracten conjuntament en una Xarxa Neural Convolucional (CNN). Atès que la segmentació semàntica i l'estimació de profunditat són dues tasques altament correlacionades, l'aproximació a les mateixes pot ser mútuament beneficiosa. En aquest cas, la informació de profunditat juntament amb l'anotació de segmentació en les dades d'entrenament ajuda a definir millor l'objectiu del procés d'entrenament del classificador, en comptes d'alimentar el sistema cegament amb un canal d'entrada addicional. Dissenyem una nova arquitectura híbrida CNN investigant els atributs comuns, així com la distinció per a l'estimació de profunditat i la segmentació semàntica. L'arquitectura proposada es prova i es compara amb l'estat de l'art en diferents conjunts de dades. Encara que s'obtenen resultats excel·lents en la segmentació semàntica, les limitacions d'aquests enfocaments també són evidents. La segmentació semàntica es recolza fortament en la semàntica predefinida i una gran quantitat de dades anotades, que potser no estaran disponibles en aplicacions més generals. D'altra banda, la segmentació d'imatge clàssica aborda la tasca de segmentació d'una manera més general. Però els enfocaments clàssics gairebé no aconsegueixen la segmentació a nivell d'objectes a causa de la manca de coneixements de nivell superior. Així, en la segona part d'aquesta tesi, ens centrem en un problema de segmentació d'instàncies genèric basat en RGB-D, on la informació temporal està disponible a partir del vídeo RGB-D, mentre que no es proporciona informació semàntica. Presentem un nou enfocament genèric de segmentació per a vídeos de núvols de punts 3D explotant a fons la geometria explícita i les correspondències temporals en RGB-D. L'enfocament proposat es valida i es compara amb enfocaments de segmentació genèrica de l'estat de l'art en diferents conjunts de dades. Finalment, en la tercera part d'aquesta tesi, presentem un mètode que combina els avantatges tant en la segmentació semàntica com en la segmentació genèrica, on descobrim instàncies de l'objecte utilitzant l'enfocament genèric i les modelem mitjançant l'aprenentatge dels pocs exemples descoberts aplicant l'enfocament de segmentació semàntica. Per fer-ho, utilitzem la tècnica d'aprenentatge d'un tir, que realitza la transferència de coneixement d'un model entrenat de forma genèrica a un model d'instància específic. Els models apresos d'instància generen funcions robustes per distingir diferents instàncies, que alimenten la segmentació genèrica de segmentació per a la seva millora. L'enfocament es valida amb experiments realitzats en un conjunt de dades acuradament seleccionat

LAReferencia - Red Federada de Repositorios Institucionales de Publicaciones Científicas Latinoamericanas

UPCommons. Portal del coneixement obert de la UPC

Tesis Doctorals en Xarxa

Semantic and generic object segmentation for scene analysis using RGB-D Data

Author: Lin Xiao
Publication venue: Universitat Politècnica de Catalunya
Publication date: 20/07/2018
Field of study

UPCommons. Portal del coneixement obert de la UPC

The correlation between vehicle vertical dynamics and deep learning-based visual target state estimation:A sensitivity study

Author: Lin
Redmon
Redmon
Ren
Simonyan
Stratis Kanarachos
Yannik Weber
Publication venue: 'MDPI AG'
Publication date: 08/11/2019
Field of study

Automated vehicles will provide greater transport convenience and interconnectivity, increase mobility options to young and elderly people, and reduce traffic congestion and emissions. However, the largest obstacle towards the deployment of automated vehicles on public roads is their safety evaluation and validation. Undeniably, the role of cameras and Artificial Intelligence-based (AI) vision is vital in the perception of the driving environment and road safety. Although a significant number of studies on the detection and tracking of vehicles have been conducted, none of them focused on the role of vertical vehicle dynamics. For the first time, this paper analyzes and discusses the influence of road anomalies and vehicle suspension on the performance of detecting and tracking driving objects. To this end, we conducted an extensive road field study and validated a computational tool for performing the assessment using simulations. A parametric study revealed the cases where AI-based vision underperforms and may significantly degrade the safety performance of AV

Crossref

ZENODO

Coventry University Pure Portal

NEUROSURGERY ENTHUSIASTIC WOMEN SOCIETY

Analyzing and controlling large nanosystems with physics-trained neural networks

Author: Stielow Thomas Gerhard Heinrich (gnd: 1262432855)
Publication venue: Universität Rostock Rostock
Publication date
Field of study

In dieser Arbeit wird untersucht, wie Neuronale Netze genutzt werden können, um die Auswertung von Experimenten durch Minimierung des Simulationsaufwandes beschleunigen zu können. Für die Rekonstruktion von Silber-Nanoclustern aus Einzelschuss-Weitwinkel-Streubildern können diese bereits aus kleinen Datenätzen allgemeine Rekonstruktionsregeln ableiten und ermöglichen durch direktes Training auf der Streuphysik unerreichte Detailtiefen. Für Giant-Dipole-Zustände von Rydbergexzitonen in Kupferoxydul wird mittels Deep Reinforcement Learning ein Anregungsschema aus Simulationen hergeleitet.This thesis investigates the possible application of neural networks in accelerating the evaluation of physical experiments while minimizing the required simulation effort. Neural networks are capable of inferring universal reconstruction rules for reconstructing silver nanoclusters from single wide-angle scattering patterns from a small set of simulated data and when trained directly on scattering theory reaching unmatched accuracy. A dynamic excitation for giant dipole states of Rydberg excitons in cuprous oxide is derived through deep reinforcement learning interacting and simulation data

Rostocker Dokumentenserver

Development of a real-time classifier for the identification of the Sit-To-Stand motion pattern

Author: Job Mirko
Publication venue: Universit\ue0 degli studi di Genova
Publication date: 15/07/2021
Field of study

The Sit-to-Stand (STS) movement has significant importance in clinical practice, since it is an indicator of lower limb functionality. As an optimal trade-off between costs and accuracy, accelerometers have recently been used to synchronously recognise the STS transition in various Human Activity Recognition-based tasks. However, beyond the mere identification of the entire action, a major challenge remains the recognition of clinically relevant phases inside the STS motion pattern, due to the intrinsic variability of the movement. This work presents the development process of a deep-learning model aimed at recognising specific clinical valid phases in the STS, relying on a pool of 39 young and healthy participants performing the task under self-paced (SP) and controlled speed (CT). The movements were registered using a total of 6 inertial sensors, and the accelerometric data was labelised into four sequential STS phases according to the Ground Reaction Force profiles acquired through a force plate. The optimised architecture combined convolutional and recurrent neural networks into a hybrid approach and was able to correctly identify the four STS phases, both under SP and CT movements, relying on the single sensor placed on the chest. The overall accuracy estimate (median [95% confidence intervals]) for the hybrid architecture was 96.09 [95.37 - 96.56] in SP trials and 95.74 [95.39 \u2013 96.21] in CT trials. Moreover, the prediction delays ( 4533 ms) were compatible with the temporal characteristics of the dataset, sampled at 10 Hz (100 ms). These results support the implementation of the proposed model in the development of digital rehabilitation solutions able to synchronously recognise the STS movement pattern, with the aim of effectively evaluate and correct its execution

Archivio istituzionale della ricerca - Università di Genova

Gaze-Based Human-Robot Interaction by the Brunswick Model

Author: A Vinciarelli
Antonia F. de C. Hamilton
B Sadrfaridpour
C Breazeal
C Breazeal
C Breazeal
C Goodwin
C Yu
D McColl
E Brunswik
G Skantze
GJM Kruijff
H Admoni
KM Lee
M Argyle
MA Yousuf
ML Walters
N Ambady
R Hartson
S Lemaignan
Publication venue
Publication date: 01/01/2019
Field of study

We present a new paradigm for human-robot interaction based on social signal processing, and in particular on the Brunswick model. Originally, the Brunswick model copes with face-to-face dyadic interaction, assuming that the interactants are communicating through a continuous exchange of non verbal social signals, in addition to the spoken messages. Social signals have to be interpreted, thanks to a proper recognition phase that considers visual and audio information. The Brunswick model allows to quantitatively evaluate the quality of the interaction using statistical tools which measure how effective is the recognition phase. In this paper we cast this theory when one of the interactants is a robot; in this case, the recognition phase performed by the robot and the human have to be revised w.r.t. the original model. The model is applied to Berrick, a recent open-source low-cost robotic head platform, where the gazing is the social signal to be considered

Crossref

Catalogo dei prodotti della ricerca

Open Access Repository