Search CORE

3,279 research outputs found

Rapid Visual Categorization is not Guided by Early Salience-Based Selection

Author: Kotseruba Iuliia
Tsotsos John K.
Wloka Calden
Publication venue: 'Public Library of Science (PLoS)'
Publication date: 01/01/2019
Field of study

The current dominant visual processing paradigm in both human and machine research is the feedforward, layered hierarchy of neural-like processing elements. Within this paradigm, visual saliency is seen by many to have a specific role, namely that of early selection. Early selection is thought to enable very fast visual performance by limiting processing to only the most salient candidate portions of an image. This strategy has led to a plethora of saliency algorithms that have indeed improved processing time efficiency in machine algorithms, which in turn have strengthened the suggestion that human vision also employs a similar early selection strategy. However, at least one set of critical tests of this idea has never been performed with respect to the role of early selection in human vision. How would the best of the current saliency models perform on the stimuli used by experimentalists who first provided evidence for this visual processing paradigm? Would the algorithms really provide correct candidate sub-images to enable fast categorization on those same images? Do humans really need this early selection for their impressive performance? Here, we report on a new series of tests of these questions whose results suggest that it is quite unlikely that such an early selection process has any role in human rapid visual categorization.Comment: 22 pages, 9 figure

arXiv.org e-Print Archive

Directory of Open Access Journals

Objects predict fixations better than early saliency

Author: Einhäuser Wolfgang
Perona Pietro
Spain Merrielle
Publication venue: 'Association for Research in Vision and Ophthalmology (ARVO)'
Publication date: 20/11/2008
Field of study

Humans move their eyes while looking at scenes and pictures. Eye movements correlate with shifts in attention and are thought to be a consequence of optimal resource allocation for high-level tasks such as visual recognition. Models of attention, such as “saliency maps,” are often built on the assumption that “early” features (color, contrast, orientation, motion, and so forth) drive attention directly. We explore an alternative hypothesis: Observers attend to “interesting” objects. To test this hypothesis, we measure the eye position of human observers while they inspect photographs of common natural scenes. Our observers perform different tasks: artistic evaluation, analysis of content, and search. Immediately after each presentation, our observers are asked to name objects they saw. Weighted with recall frequency, these objects predict fixations in individual images better than early saliency, irrespective of task. Also, saliency combined with object positions predicts which objects are frequently named. This suggests that early saliency has only an indirect effect on attention, acting through recognized objects. Consequently, rather than treating attention as mere preprocessing step for object recognition, models of both need to be integrated

Caltech Authors

Contextual cropping and scaling of TV productions

Author: A Treisman
DA Forsyth
DL Ruderman
Gerhard Stoll
Joerg Deigmoeller
L Itti
L Sachs
L-Q Chen
M Knee
Norbert Just
O Meur Le
R Mohan
Takebumi Itagaki
W-H Cheng
WY Lum
X Hou
Z Zhang
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 11/05/2011
Field of study

This is the author's accepted manuscript. The final publication is available at Springer via http://dx.doi.org/10.1007/s11042-011-0804-3. Copyright @ Springer Science+Business Media, LLC 2011.In this paper, an application is presented which automatically adapts SDTV (Standard Definition Television) sports productions to smaller displays through intelligent cropping and scaling. It crops regions of interest of sports productions based on a smart combination of production metadata and systematic video analysis methods. This approach allows a context-based composition of cropped images. It provides a differentiation between the original SD version of the production and the processed one adapted to the requirements for mobile TV. The system has been comprehensively evaluated by comparing the outcome of the proposed method with manually and statically cropped versions, as well as with non-cropped versions. Envisaged is the integration of the tool in post-production and live workflows

Crossref

Brunel University Research Archive

Identifying and recognizing noticeable sounds from physical measurements and their effect on soundscape

Author: Boes Michiel
Botteldooren Dick
De Coensel Bert
Domitrović Hrvoje
Filipan Karlo
Publication venue
Publication date: 01/01/2015
Field of study

Ghent University Academic Bibliography