Search CORE

6,985 research outputs found

Video Upright Adjustment and Stabilization

Author: Jucheol Won
Publication venue: Daegu
Publication date: 20/01/2020
Field of study

Upright adjustment, Video stabilization, Camera pathWe propose a novel video upright adjustment method that can reliably correct slanted video contents that are often found in casual videos. Our approach combines deep learning and Bayesian inference to estimate accurate rotation angles from video frames. We train a convolutional neural network to obtain initial estimates of the rotation angles of input video frames. The initial estimates from the network are temporally inconsistent and inaccurate. To resolve this, we use Bayesian inference. We analyze estimation errors of the network, and derive an error model. We then use the error model to formulate video upright adjustment as a maximum a posteriori problem where we estimate consistent rotation angles from the initial estimates, while respecting relative rotations between consecutive frames. Finally, we propose a joint approach to video stabilization and upright adjustment, which minimizes information loss caused by separately handling stabilization and upright adjustment. Experimental results show that our video upright adjustment method can effectively correct slanted video contents, and its combination with video stabilization can achieve visually pleasing results from shaky and slanted videos.openI. INTRODUCTION 1.1. Related work II. ROTATION ESTIMATION NETWORK III. ERROR ANALYSIS IV. VIDEO UPRIGHT ADJUSTMENT 4.1. Initial angle estimation 4.2. Robust angle estimation 4.3. Optimization 4.4. Warping V. JOINT UPRIGHT ADJUSTMENT AND STABILIZATION 5.1. Bundled camera paths for video stabilization 5.2. Joint approach VI. EXPERIMENTS VII. CONCLUSION ReferencesCNN)을 훈련시킨다. 신경망의 초기 추정치는 완전히 정확하지 않으며 시간적으로도 일관되지 않는다. 이를 해결하기 위해 베이지안 인퍼런스를 사용한다. 본 논문은 신경망의 추정 오류를 분석하고 오류 모델을 도출한다. 그런 다음 오류 모델을 사용하여 연속 프레임 간의 상대 회전 각도(Relative rotation angle)를 반영하면서 초기 추정치로부터 시간적으로 일관된 회전 각도를 추정하는 최대 사후 문제(Maximum a posteriori problem)로 동영상 수평 보정을 공식화한다. 마지막으로, 동영상 수평 보정 및 동영상 안정화(Video stabilization)에 대한 동시 접근 방법을 제안하여 수평 보정과 안정화를 별도로 수행할 때 발생하는 공간 정보 손실과 연산량을 최소화하며 안정화의 성능을 최대화한다. 실험 결과에 따르면 동영상 수평 보정으로 기울어진 동영상을 효과적으로 보정할 수 있으며 동영상 안정화 방법과 결합하여 흔들리고 기울어진 동영상으로부터 시각적으로 만족스러운 새로운 동영상을 획득할 수 있다.본 논문은 일반인들이 촬영한 동영상에서 흔히 발생하는 문제인 기울어짐을 제거하여 수평이 올바른 동영상을 획득할 수 있게 하는 동영상 수평 보정(Video upright adjustment) 방법을 제안한다. 본 논문의 접근 방식은 딥 러닝(Deep learning)과 베이지안 인퍼런스(Bayesian inference)를 결합하여 동영상 프레임(Frame)에서 정확한 각도를 추정한다. 먼저 입력 동영상 프레임의 회전 각도의 초기 추정치를 얻기 위해 회선 신경망(Convolutional neural networkMasterdCollectio

DGIST Library Institutional Repository

Transitioning360: Content-aware NFoV Virtual Camera Paths for 360° Video Playback

Author: Hu Shi-Min
Li Yi-Jun
Richardt Christian
Wang Miao
Zhang Wen-Xuan
Publication venue: 'Institute of Electrical and Electronics Engineers (IEEE)'
Publication date: 31/12/2020
Field of study

OPUS

The highD Dataset: A Drone Dataset of Naturalistic Vehicle Trajectories on German Highways for Validation of Highly Automated Driving Systems

Author: Bock Julian
Eckstein Lutz
Kloeker Laurent
Krajewski Robert
Publication venue
Publication date: 01/01/2018
Field of study

Scenario-based testing for the safety validation of highly automated vehicles is a promising approach that is being examined in research and industry. This approach heavily relies on data from real-world scenarios to derive the necessary scenario information for testing. Measurement data should be collected at a reasonable effort, contain naturalistic behavior of road users and include all data relevant for a description of the identified scenarios in sufficient quality. However, the current measurement methods fail to meet at least one of the requirements. Thus, we propose a novel method to measure data from an aerial perspective for scenario-based validation fulfilling the mentioned requirements. Furthermore, we provide a large-scale naturalistic vehicle trajectory dataset from German highways called highD. We evaluate the data in terms of quantity, variety and contained scenarios. Our dataset consists of 16.5 hours of measurements from six locations with 110 000 vehicles, a total driven distance of 45 000 km and 5600 recorded complete lane changes. The highD dataset is available online at: http://www.highD-dataset.comComment: IEEE International Conference on Intelligent Transportation Systems (ITSC) 201

arXiv.org e-Print Archive

Crossref

Publikationsserver der RWTH Aachen University

Long-Term Visual Object Tracking Benchmark

Author: AW Smeulders
B Babenko
C Vondrick
D Held
H Grabner
H Li
J Zhang
Jack Valmadre
JF Henriques
JF Henriques
M Danelljan
M Kristan
M Kumar
M Mueller
P Liang
WL Lu
Y Hua
Y Li
Y Wu
Z Kalal
Publication venue
Publication date: 01/01/2019
Field of study

We propose a new long video dataset (called Track Long and Prosper - TLP) and benchmark for single object tracking. The dataset consists of 50 HD videos from real world scenarios, encompassing a duration of over 400 minutes (676K frames), making it more than 20 folds larger in average duration per sequence and more than 8 folds larger in terms of total covered duration, as compared to existing generic datasets for visual tracking. The proposed dataset paves a way to suitably assess long term tracking performance and train better deep learning architectures (avoiding/reducing augmentation, which may not reflect real world behaviour). We benchmark the dataset on 17 state of the art trackers and rank them according to tracking accuracy and run time speeds. We further present thorough qualitative and quantitative evaluation highlighting the importance of long term aspect of tracking. Our most interesting observations are (a) existing short sequence benchmarks fail to bring out the inherent differences in tracking algorithms which widen up while tracking on long sequences and (b) the accuracy of trackers abruptly drops on challenging long sequences, suggesting the potential need of research efforts in the direction of long-term tracking.Comment: ACCV 2018 (Oral

arXiv.org e-Print Archive

Crossref

VNect: Real-time 3D Human Pose Estimation with a Single RGB Camera

Author: Casas Dan
Mehta Dushyant
Rhodin Helge
Seidel Hans-Peter
Shafiei Mohammad
Sotnychenko Oleksandr
Sridhar Srinath
Theobalt Christian
Xu Weipeng
Publication venue: 'Association for Computing Machinery (ACM)'
Publication date: 01/01/2017
Field of study

We present the first real-time method to capture the full global 3D skeletal pose of a human in a stable, temporally consistent manner using a single RGB camera. Our method combines a new convolutional neural network (CNN) based pose regressor with kinematic skeleton fitting. Our novel fully-convolutional pose formulation regresses 2D and 3D joint positions jointly in real time and does not require tightly cropped input frames. A real-time kinematic skeleton fitting method uses the CNN output to yield temporally stable 3D global pose reconstructions on the basis of a coherent kinematic skeleton. This makes our approach the first monocular RGB method usable in real-time applications such as 3D character control---thus far, the only monocular methods for such applications employed specialized RGB-D cameras. Our method's accuracy is quantitatively on par with the best offline 3D monocular RGB pose estimation methods. Our results are qualitatively comparable to, and sometimes better than, results from monocular RGB-D approaches, such as the Kinect. However, we show that our approach is more broadly applicable than RGB-D solutions, i.e. it works for outdoor scenes, community videos, and low quality commodity RGB cameras.Comment: Accepted to SIGGRAPH 201

arXiv.org e-Print Archive

MPG.PuRe