2,347 research outputs found
On Compositional Hierarchical Models for holistic Lane and Road Perception in Intelligent Vehicles
This work is a contribution to the vision based perception of multi lane roads of urban intersections. Given multiple input features the proposed probabilistic hierarchical model infers the lane structure as well as the location of stoplines and the turn directions of individual lanes. Thereby, it expresses prior expectations on the road topology using weak probabilistic constraints which allows for the detection of parallel lanes as well as splitting and merging lanes
Hierarchische Modelle für das visuelle Erkennen und Lernen von Objekten, Szenen und Aktivitäten
In many computer vision applications, objects have to be learned and recognized in images or image sequences. Most of these objects have a hierarchical structure.For example, 3d objects can be decomposed into object parts, and object parts, in turn, into geometric primitives. Furthermore, scenes are composed of objects. And also activities or behaviors can be divided hierarchically into actions, these into individual movements, etc. Hierarchical models are therefore ideally suited for the representation of a wide range of objects used in applications such as object recognition, human pose estimation, or activity recognition.
In this work new probabilistic hierarchical models are presented that allow an efficient representation of multiple objects of different categories, scales, rotations, and views. The idea is to exploit similarities between objects, object parts or actions and movements in order to share calculations and avoid redundant information. We will introduce online and offline learning methods, which enable to create efficient hierarchies based on small or large training datasets, in which poses or articulated structures are given by instances. Furthermore, we present inference approaches for fast and robust detection. These new approaches combine the idea of compositional and similarity hierarchies and overcome limitations of previous methods. They will be used in an unified hierarchical framework spatially for object recognition as well as spatiotemporally for activity recognition.
The unified generic hierarchical framework allows us to apply the proposed models in different projects. Besides classical object recognition it is used for detection of human poses in a project for gait analysis. The activity detection is used in a project for the design of environments for ageing, to identify activities and behavior patterns in smart homes. In a project for parking spot detection using an intelligent vehicle, the proposed approaches are used to hierarchically model the environment of the vehicle for an efficient and robust interpretation of the scene in real-time.In zahlreichen Computer Vision Anwendungen müssen Objekte in einzelnen Bildern oder Bildsequenzen erlernt und erkannt werden. Viele dieser Objekte sind hierarchisch aufgebaut.So lassen sich 3d Objekte in Objektteile zerlegen und Objektteile wiederum in geometrische Grundkörper. Und auch Aktivitäten oder Verhaltensmuster lassen sich hierarchisch in einzelne Aktionen aufteilen, diese wiederum in einzelne Bewegungen usw. Für die Repräsentation sind hierarchische Modelle dementsprechend gut geeignet.
In dieser Arbeit werden neue probabilistische hierarchische Modelle vorgestellt, die es ermöglichen auch mehrere Objekte verschiedener Kategorien, Skalierungen, Rotationen und aus verschiedenen Blickrichtungen effizient zu repräsentieren. Eine Idee ist hierbei, Ähnlichkeiten unter Objekten, Objektteilen oder auch Aktionen und Bewegungen zu nutzen, um redundante Informationen und Mehrfachberechnungen zu vermeiden. In der Arbeit werden online und offline Lernverfahren vorgestellt, die es ermöglichen, effiziente Hierarchien auf Basis von kleinen oder großen Trainingsdatensätzen zu erstellen, in denen Posen und bewegliche Strukturen durch Beispiele gegeben sind. Des Weiteren werden Inferenzansätze zur schnellen und robusten Detektion vorgestellt. Diese werden innerhalb eines einheitlichen hierarchischen Frameworks sowohl räumlich zur Objekterkennung als auch raumzeitlich zur Aktivitätenerkennung verwendet.
Das einheitliche Framework ermöglicht die Anwendung des vorgestellten Modells innerhalb verschiedener Projekte. Neben der klassischen Objekterkennung wird es zur Erkennung von menschlichen Posen in einem Projekt zur Ganganalyse verwendet. Die Aktivitätenerkennung wird in einem Projekt zur Gestaltung altersgerechter Lebenswelten genutzt, um in intelligenten Wohnräumen Aktivitäten und Verhaltensmuster von Bewohnern zu erkennen. Im Rahmen eines Projektes zur Parklückenvermessung mithilfe eines intelligenten Fahrzeuges werden die vorgestellten Ansätze verwendet, um das Umfeld des Fahrzeuges hierarchisch zu modellieren und dadurch das Szenenverstehen zu ermöglichen
Holistic Temporal Situation Interpretation for Traffic Participant Prediction
For a profound understanding of traffic situations including a prediction of traf-
fic participants’ future motion, behaviors and routes it is crucial to incorporate all
available environmental observations. The presence of sensor noise and depen-
dency uncertainties, the variety of available sensor data, the complexity of large
traffic scenes and the large number of different estimation tasks with diverging
requirements require a general method that gives a robust foundation for the de-
velopment of estimation applications.
In this work, a general description language, called Object-Oriented Factor Graph
Modeling Language (OOFGML), is proposed, that unifies formulation of esti-
mation tasks from the application-oriented problem description via the choice
of variable and probability distribution representation through to the inference
method definition in implementation. The different language properties are dis-
cussed theoretically using abstract examples.
The derivation of explicit application examples is shown for the automated driv-
ing domain. A domain-specific ontology is defined which forms the basis for
four exemplary applications covering the broad spectrum of estimation tasks in
this domain: Basic temporal filtering, ego vehicle localization using advanced
interpretations of perceived objects, road layout perception utilizing inter-object
dependencies and finally highly integrated route, behavior and motion estima-
tion to predict traffic participant’s future actions. All applications are evaluated
as proof of concept and provide an example of how their class of estimation tasks
can be represented using the proposed language. The language serves as a com-
mon basis and opens a new field for further research towards holistic solutions
for automated driving
Building Machines That Learn and Think Like People
Recent progress in artificial intelligence (AI) has renewed interest in
building systems that learn and think like people. Many advances have come from
using deep neural networks trained end-to-end in tasks such as object
recognition, video games, and board games, achieving performance that equals or
even beats humans in some respects. Despite their biological inspiration and
performance achievements, these systems differ from human intelligence in
crucial ways. We review progress in cognitive science suggesting that truly
human-like learning and thinking machines will have to reach beyond current
engineering trends in both what they learn, and how they learn it.
Specifically, we argue that these machines should (a) build causal models of
the world that support explanation and understanding, rather than merely
solving pattern recognition problems; (b) ground learning in intuitive theories
of physics and psychology, to support and enrich the knowledge that is learned;
and (c) harness compositionality and learning-to-learn to rapidly acquire and
generalize knowledge to new tasks and situations. We suggest concrete
challenges and promising routes towards these goals that can combine the
strengths of recent neural network advances with more structured cognitive
models.Comment: In press at Behavioral and Brain Sciences. Open call for commentary
proposals (until Nov. 22, 2016).
https://www.cambridge.org/core/journals/behavioral-and-brain-sciences/information/calls-for-commentary/open-calls-for-commentar
A LiDAR Point Cloud Generator: from a Virtual World to Autonomous Driving
3D LiDAR scanners are playing an increasingly important role in autonomous
driving as they can generate depth information of the environment. However,
creating large 3D LiDAR point cloud datasets with point-level labels requires a
significant amount of manual annotation. This jeopardizes the efficient
development of supervised deep learning algorithms which are often data-hungry.
We present a framework to rapidly create point clouds with accurate point-level
labels from a computer game. The framework supports data collection from both
auto-driving scenes and user-configured scenes. Point clouds from auto-driving
scenes can be used as training data for deep learning algorithms, while point
clouds from user-configured scenes can be used to systematically test the
vulnerability of a neural network, and use the falsifying examples to make the
neural network more robust through retraining. In addition, the scene images
can be captured simultaneously in order for sensor fusion tasks, with a method
proposed to do automatic calibration between the point clouds and captured
scene images. We show a significant improvement in accuracy (+9%) in point
cloud segmentation by augmenting the training dataset with the generated
synthesized data. Our experiments also show by testing and retraining the
network using point clouds from user-configured scenes, the weakness/blind
spots of the neural network can be fixed
Predictive World Models from Real-World Partial Observations
Cognitive scientists believe adaptable intelligent agents like humans perform
reasoning through learned causal mental simulations of agents and environments.
The problem of learning such simulations is called predictive world modeling.
Recently, reinforcement learning (RL) agents leveraging world models have
achieved SOTA performance in game environments. However, understanding how to
apply the world modeling approach in complex real-world environments relevant
to mobile robots remains an open question. In this paper, we present a
framework for learning a probabilistic predictive world model for real-world
road environments. We implement the model using a hierarchical VAE (HVAE)
capable of predicting a diverse set of fully observed plausible worlds from
accumulated sensor observations. While prior HVAE methods require complete
states as ground truth for learning, we present a novel sequential training
method to allow HVAEs to learn to predict complete states from partially
observed states only. We experimentally demonstrate accurate spatial structure
prediction of deterministic regions achieving 96.21 IoU, and close the gap to
perfect prediction by 62% for stochastic regions using the best prediction. By
extending HVAEs to cases where complete ground truth states do not exist, we
facilitate continual learning of spatial prediction as a step towards realizing
explainable and comprehensive predictive world models for real-world mobile
robotics applications. Code is available at
https://github.com/robin-karlsson0/predictive-world-models.Comment: Accepted for IEEE MOST 202
- …