Search CORE

10,603 research outputs found

Visual to Sound: Generating Natural Sound for Videos in the Wild

Author: Berg Tamara L.
Bui Trung
Fang Chen
Wang Zhaowen
Zhou Yipin
Publication venue
Publication date: 01/06/2018
Field of study

As two of the five traditional human senses (sight, hearing, taste, smell, and touch), vision and sound are basic sources through which humans understand the world. Often correlated during natural events, these two modalities combine to jointly affect human perception. In this paper, we pose the task of generating sound given visual input. Such capabilities could help enable applications in virtual reality (generating sound for virtual scenes automatically) or provide additional accessibility to images or videos for people with visual impairments. As a first step in this direction, we apply learning-based methods to generate raw waveform samples given input video frames. We evaluate our models on a dataset of videos containing a variety of sounds (such as ambient sounds and sounds from people/animals). Our experiments show that the generated sounds are fairly realistic and have good temporal synchronization with the visual inputs.Comment: Project page: http://bvision11.cs.unc.edu/bigpen/yipin/visual2sound_webpage/visual2sound.htm

arXiv.org e-Print Archive

Crossref

The Fine Art of Commercial Freedom: British Music Videos and Film Culture

Author: Caston Emily
Publication venue
Publication date: 01/02/2014
Field of study

An outline and analysis of the British music video industry and its impact on film culture in the 1990s and 2000s

UAL Research Online

UWL Repository

Hierarchical Cross-Modal Talking Face Generationwith Dynamic Pixel-Wise Loss

Author: Chen Lele
Duan Zhiyao
Maddox Ross K.
Xu Chenliang
Publication venue
Publication date: 09/05/2019
Field of study

We devise a cascade GAN approach to generate talking face video, which is robust to different face shapes, view angles, facial characteristics, and noisy audio conditions. Instead of learning a direct mapping from audio to video frames, we propose first to transfer audio to high-level structure, i.e., the facial landmarks, and then to generate video frames conditioned on the landmarks. Compared to a direct audio-to-image approach, our cascade approach avoids fitting spurious correlations between audiovisual signals that are irrelevant to the speech content. We, humans, are sensitive to temporal discontinuities and subtle artifacts in video. To avoid those pixel jittering problems and to enforce the network to focus on audiovisual-correlated regions, we propose a novel dynamically adjustable pixel-wise loss with an attention mechanism. Furthermore, to generate a sharper image with well-synchronized facial movements, we propose a novel regression-based discriminator structure, which considers sequence-level information along with frame-level information. Thoughtful experiments on several datasets and real-world samples demonstrate significantly better results obtained by our method than the state-of-the-art methods in both quantitative and qualitative comparisons

arXiv.org e-Print Archive

Crossref

The Celtic Tiger ‘Unplugged’: DV realism, liveness, and sonic authenticity in Once (2007)

Author: Johnston Nessa
Publication venue: 'Intellect'
Publication date: 01/04/2014
Field of study

Crossref

Edge Hill University Research Information Repository

RGB-D datasets using microsoft kinect or similar sensors: a survey

Author: Galili
Guan
Hu
Kolner
Mulvad
Nakazawa
Palushani
Palushani
Publication venue: Springer
Publication date: 01/01/2015
Field of study

RGB-D data has turned out to be a very useful representation of an indoor scene for solving fundamental computer vision problems. It takes the advantages of the color image that provides appearance information of an object and also the depth image that is immune to the variations in color, illumination, rotation angle and scale. With the invention of the low-cost Microsoft Kinect sensor, which was initially used for gaming and later became a popular device for computer vision, high quality RGB-D data can be acquired easily. In recent years, more and more RGB-D image/video datasets dedicated to various applications have become available, which are of great importance to benchmark the state-of-the-art. In this paper, we systematically survey popular RGB-D datasets for different applications including object recognition, scene classification, hand gesture recognition, 3D-simultaneous localization and mapping, and pose estimation. We provide the insights into the characteristics of each important dataset, and compare the popularity and the difficulty of those datasets. Overall, the main goal of this survey is to give a comprehensive description about the available RGB-D datasets and thus to guide researchers in the selection of suitable datasets for evaluating their algorithms

Northumbria Research Link

Crossref

Springer - Publisher Connector

Online Research Database In Technology

Smart Concerts: Orchestras in the Age of Edutainment

Author: Alan S. Brown
Publication venue: John S. and James L. Knight Foundation
Publication date: 01/01/2005
Field of study

Provides a summary of the recent shift in concert programming, and discusses four strategies for enhancing the concert experience: contextual programming, dramatization of music, visual enhancements, and embedded interpretation

IssueLab