11,866 research outputs found

    End-to-end Driving via Conditional Imitation Learning

    Get PDF
    Deep networks trained on demonstrations of human driving have learned to follow roads and avoid obstacles. However, driving policies trained via imitation learning cannot be controlled at test time. A vehicle trained end-to-end to imitate an expert cannot be guided to take a specific turn at an upcoming intersection. This limits the utility of such systems. We propose to condition imitation learning on high-level command input. At test time, the learned driving policy functions as a chauffeur that handles sensorimotor coordination but continues to respond to navigational commands. We evaluate different architectures for conditional imitation learning in vision-based driving. We conduct experiments in realistic three-dimensional simulations of urban driving and on a 1/5 scale robotic truck that is trained to drive in a residential area. Both systems drive based on visual input yet remain responsive to high-level navigational commands. The supplementary video can be viewed at https://youtu.be/cFtnflNe5fMComment: Published at the International Conference on Robotics and Automation (ICRA), 201

    Towards End-to-End Acoustic Localization using Deep Learning: from Audio Signal to Source Position Coordinates

    Full text link
    This paper presents a novel approach for indoor acoustic source localization using microphone arrays and based on a Convolutional Neural Network (CNN). The proposed solution is, to the best of our knowledge, the first published work in which the CNN is designed to directly estimate the three dimensional position of an acoustic source, using the raw audio signal as the input information avoiding the use of hand crafted audio features. Given the limited amount of available localization data, we propose in this paper a training strategy based on two steps. We first train our network using semi-synthetic data, generated from close talk speech recordings, and where we simulate the time delays and distortion suffered in the signal that propagates from the source to the array of microphones. We then fine tune this network using a small amount of real data. Our experimental results show that this strategy is able to produce networks that significantly improve existing localization methods based on \textit{SRP-PHAT} strategies. In addition, our experiments show that our CNN method exhibits better resistance against varying gender of the speaker and different window sizes compared with the other methods.Comment: 18 pages, 3 figures, 8 table

    Which tone-mapping operator is the best? A comparative study of perceptual quality

    Get PDF
    Altres ajuts: CERCA Programme/Generalitat de CatalunyaPublicat sota la llicència Open Access Publishing Agreement, específica d'Optica Publishing Group https://opg.optica.org/submit/review/pdf/CopyrightTransferOpenAccessAgreement-2022-06-27.pdfTone-mapping operators (TMOs) are designed to generate perceptually similar low-dynamic-range images from high-dynamic-range ones. We studied the performance of 15 TMOs in two psychophysical experiments where observers compared the digitally generated tone-mapped images to their corresponding physical scenes. All experiments were performed in a controlled environment, and the setups were designed to emphasize different image properties: in the first experiment we evaluated the local relationships among intensity levels, and in the second one we evaluated global visual appearance among physical scenes and tone-mapped images, which were presented side by side. We ranked the TMOs according to how well they reproduced the results obtained in the physical scene. Our results show that ranking position clearly depends on the adopted evaluation criteria, which implies that, in general, these tone-mapping algorithms consider either local or global image attributes but rarely both. Regarding the question of which TMO is the best, KimKautz ["Consistent tone reproduction," in Proceedings of Computer Graphics and Imaging (2008)] and Krawczyk ["Lightness perception in tone reproduction for high dynamic range images," in Proceedings of Eurographics (2005), p. 3] obtained the better results across the different experiments. We conclude that more thorough and standardized evaluation criteria are needed to study all the characteristics of TMOs, as there is ample room for improvement in future developments

    Algorithms for compression of high dynamic range images and video

    Get PDF
    The recent advances in sensor and display technologies have brought upon the High Dynamic Range (HDR) imaging capability. The modern multiple exposure HDR sensors can achieve the dynamic range of 100-120 dB and LED and OLED display devices have contrast ratios of 10^5:1 to 10^6:1. Despite the above advances in technology the image/video compression algorithms and associated hardware are yet based on Standard Dynamic Range (SDR) technology, i.e. they operate within an effective dynamic range of up to 70 dB for 8 bit gamma corrected images. Further the existing infrastructure for content distribution is also designed for SDR, which creates interoperability problems with true HDR capture and display equipment. The current solutions for the above problem include tone mapping the HDR content to fit SDR. However this approach leads to image quality associated problems, when strong dynamic range compression is applied. Even though some HDR-only solutions have been proposed in literature, they are not interoperable with current SDR infrastructure and are thus typically used in closed systems. Given the above observations a research gap was identified in the need for efficient algorithms for the compression of still images and video, which are capable of storing full dynamic range and colour gamut of HDR images and at the same time backward compatible with existing SDR infrastructure. To improve the usability of SDR content it is vital that any such algorithms should accommodate different tone mapping operators, including those that are spatially non-uniform. In the course of the research presented in this thesis a novel two layer CODEC architecture is introduced for both HDR image and video coding. Further a universal and computationally efficient approximation of the tone mapping operator is developed and presented. It is shown that the use of perceptually uniform colourspaces for internal representation of pixel data enables improved compression efficiency of the algorithms. Further proposed novel approaches to the compression of metadata for the tone mapping operator is shown to improve compression performance for low bitrate video content. Multiple compression algorithms are designed, implemented and compared and quality-complexity trade-offs are identified. Finally practical aspects of implementing the developed algorithms are explored by automating the design space exploration flow and integrating the high level systems design framework with domain specific tools for synthesis and simulation of multiprocessor systems. The directions for further work are also presented

    Highlights Analysis System (HAnS) for low dynamic range to high dynamic range conversion of cinematic low dynamic range content

    Get PDF
    We propose a novel and efficient algorithm for detection of specular reflections and light sources (highlights) in cinematic content. The detection of highlights is important for reconstructing them properly in the conversion of the low dynamic range (LDR) to high dynamic range (HDR) content. Highlights are often difficult to be distinguished from bright diffuse surfaces, due to their brightness being reduced in the conventional LDR content production. Moreover, the cinematic LDR content is subject to the artistic use of effects that change the apparent brightness of certain image regions (e.g. limiting depth of field, grading, complex multi-lighting setup, etc.). To ensure the robustness of highlights detection to these effects, the proposed algorithm goes beyond considering only absolute brightness and considers five different features. These features are: the size of the highlight relative to the size of the surrounding image structures, the relative contrast in the surrounding of the highlight, its absolute brightness expressed through the luminance (luma feature), through the saturation in the color space (maxRGB feature) and through the saturation in white (minRGB feature). We evaluate the algorithm on two different image data-sets. The first one is a publicly available LDR image data-set without cinematic content, which allows comparison to the broader State of the art. Additionally, for the evaluation on cinematic content, we create an image data-set consisted of manually annotated cinematic frames and real-world images. For the purpose of demonstrating the proposed highlights detection algorithm in a complete LDR-to-HDR conversion pipeline, we additionally propose a simple inverse-tone-mapping algorithm. The experimental analysis shows that the proposed approach outperforms conventional highlights detection algorithms on both image data-sets, achieves high quality reconstruction of the HDR content and is suited for use in LDR-to-HDR conversion
    corecore