5,213 research outputs found
A bayesian approach to simultaneously recover camera pose and non-rigid shape from monocular images
© . This manuscript version is made available under the CC-BY-NC-ND 4.0 license http://creativecommons.org/licenses/by-nc-nd/4.0/In this paper we bring the tools of the Simultaneous Localization and Map Building (SLAM) problem from a rigid to a deformable domain and use them to simultaneously recover the 3D shape of non-rigid surfaces and the sequence of poses of a moving camera. Under the assumption that the surface shape may be represented as a weighted sum of deformation modes, we show that the problem of estimating the modal weights along with the camera poses, can be probabilistically formulated as a maximum a posteriori estimate and solved using an iterative least squares optimization. In addition, the probabilistic formulation we propose is very general and allows introducing different constraints without requiring any extra complexity. As a proof of concept, we show that local inextensibility constraints that prevent the surface from stretching can be easily integrated.
An extensive evaluation on synthetic and real data, demonstrates that our method has several advantages over current non-rigid shape from motion approaches. In particular, we show that our solution is robust to large amounts of noise and outliers and that it does not need to track points over the whole sequence nor to use an initialization close from the ground truth.Peer ReviewedPostprint (author's final draft
Structure from Recurrent Motion: From Rigidity to Recurrency
This paper proposes a new method for Non-Rigid Structure-from-Motion (NRSfM)
from a long monocular video sequence observing a non-rigid object performing
recurrent and possibly repetitive dynamic action. Departing from the
traditional idea of using linear low-order or lowrank shape model for the task
of NRSfM, our method exploits the property of shape recurrency (i.e., many
deforming shapes tend to repeat themselves in time). We show that recurrency is
in fact a generalized rigidity. Based on this, we reduce NRSfM problems to
rigid ones provided that certain recurrency condition is satisfied. Given such
a reduction, standard rigid-SfM techniques are directly applicable (without any
change) to the reconstruction of non-rigid dynamic shapes. To implement this
idea as a practical approach, this paper develops efficient algorithms for
automatic recurrency detection, as well as camera view clustering via a
rigidity-check. Experiments on both simulated sequences and real data
demonstrate the effectiveness of the method. Since this paper offers a novel
perspective on rethinking structure-from-motion, we hope it will inspire other
new problems in the field.Comment: To appear in CVPR 201
Geometry-Aware Network for Non-Rigid Shape Prediction from a Single View
We propose a method for predicting the 3D shape of a deformable surface from
a single view. By contrast with previous approaches, we do not need a
pre-registered template of the surface, and our method is robust to the lack of
texture and partial occlusions. At the core of our approach is a {\it
geometry-aware} deep architecture that tackles the problem as usually done in
analytic solutions: first perform 2D detection of the mesh and then estimate a
3D shape that is geometrically consistent with the image. We train this
architecture in an end-to-end manner using a large dataset of synthetic
renderings of shapes under different levels of deformation, material
properties, textures and lighting conditions. We evaluate our approach on a
test split of this dataset and available real benchmarks, consistently
improving state-of-the-art solutions with a significantly lower computational
time.Comment: Accepted at CVPR 201
Stable Camera Motion Estimation Using Convex Programming
We study the inverse problem of estimating n locations (up to
global scale, translation and negation) in from noisy measurements of a
subset of the (unsigned) pairwise lines that connect them, that is, from noisy
measurements of for some pairs (i,j) (where the
signs are unknown). This problem is at the core of the structure from motion
(SfM) problem in computer vision, where the 's represent camera locations
in . The noiseless version of the problem, with exact line measurements,
has been considered previously under the general title of parallel rigidity
theory, mainly in order to characterize the conditions for unique realization
of locations. For noisy pairwise line measurements, current methods tend to
produce spurious solutions that are clustered around a few locations. This
sensitivity of the location estimates is a well-known problem in SfM,
especially for large, irregular collections of images.
In this paper we introduce a semidefinite programming (SDP) formulation,
specially tailored to overcome the clustering phenomenon. We further identify
the implications of parallel rigidity theory for the location estimation
problem to be well-posed, and prove exact (in the noiseless case) and stable
location recovery results. We also formulate an alternating direction method to
solve the resulting semidefinite program, and provide a distributed version of
our formulation for large numbers of locations. Specifically for the camera
location estimation problem, we formulate a pairwise line estimation method
based on robust camera orientation and subspace estimation. Lastly, we
demonstrate the utility of our algorithm through experiments on real images.Comment: 40 pages, 12 figures, 6 tables; notation and some unclear parts
updated, some typos correcte
Joint Optical Flow and Temporally Consistent Semantic Segmentation
The importance and demands of visual scene understanding have been steadily
increasing along with the active development of autonomous systems.
Consequently, there has been a large amount of research dedicated to semantic
segmentation and dense motion estimation. In this paper, we propose a method
for jointly estimating optical flow and temporally consistent semantic
segmentation, which closely connects these two problem domains and leverages
each other. Semantic segmentation provides information on plausible physical
motion to its associated pixels, and accurate pixel-level temporal
correspondences enhance the accuracy of semantic segmentation in the temporal
domain. We demonstrate the benefits of our approach on the KITTI benchmark,
where we observe performance gains for flow and segmentation. We achieve
state-of-the-art optical flow results, and outperform all published algorithms
by a large margin on challenging, but crucial dynamic objects.Comment: 14 pages, Accepted for CVRSUAD workshop at ECCV 201
In-loop Feature Tracking for Structure and Motion with Out-of-core Optimization
In this paper, a novel and approach for obtaining 3D models from video sequences captured with hand-held cameras is addressed. We define a pipeline that robustly deals with different types of sequences and acquiring devices. Our system follows a divide and conquer approach: after a frame decimation that pre-conditions the input sequence, the video is split into short-length clips. This allows to parallelize the reconstruction step which translates into a reduction in the amount of computational resources required. The short length of the clips allows an intensive search for the best solution at each step of reconstruction which robustifies the system. The process of feature tracking is embedded within the reconstruction loop for each clip as opposed to other approaches. A final registration step, merges all the processed clips to the same coordinate fram
Robust Camera Location Estimation by Convex Programming
D structure recovery from a collection of D images requires the
estimation of the camera locations and orientations, i.e. the camera motion.
For large, irregular collections of images, existing methods for the location
estimation part, which can be formulated as the inverse problem of estimating
locations in
from noisy measurements of a subset of the pairwise directions
, are
sensitive to outliers in direction measurements. In this paper, we firstly
provide a complete characterization of well-posed instances of the location
estimation problem, by presenting its relation to the existing theory of
parallel rigidity. For robust estimation of camera locations, we introduce a
two-step approach, comprised of a pairwise direction estimation method robust
to outliers in point correspondences between image pairs, and a convex program
to maintain robustness to outlier directions. In the presence of partially
corrupted measurements, we empirically demonstrate that our convex formulation
can even recover the locations exactly. Lastly, we demonstrate the utility of
our formulations through experiments on Internet photo collections.Comment: 10 pages, 6 figures, 3 table
- …