Search CORE

17 research outputs found

Desafíos del aprendizaje profundo en la visión por computador

Author: Achanccaray Díaz Pedro Marco
Arauco Canchumuni Smith Washington
Ayma Quirita Victor Hugo
Soto Vega Pedro Juan
Publication venue: 'Baishideng Publishing Group Inc.'
Publication date: 01/01/2022
Field of study

La visión por computador es un área de estudio en la inteligencia artificial que se enfoca en el desarrollo de técnicas computacionales para percibir el mundo a través de entradas visuales, como videos o imágenes. El aprendizaje profundo ha demostrado ser una técnica eficiente para el análisis e interpretación de datos visuales. Sin embargo, afronta innumerables desafíos según su aplicación en las diferentes tareas de la visión por computador. Este panel reúne un grupo de expertos en aprendizaje profundo, quienes ofrecerán información sobre su aplicación y los desafíos en sus respectivas áreas de investigación con relación a la visión por computador

Repositorio Institucional Ulima

Recognizing Vector Graphics without Rasterization

Author: Dong X
Jiang X
Li D
Liu L
Shan C
Shen Y
Publication venue: NeurIPS
Publication date: 03/07/2022
Field of study

In this paper, we consider a different data format for images: vector graphics. In contrast to raster graphics which are widely used in image recognition, vector graphics can be scaled up or down into any resolution without aliasing or information loss, due to the analytic representation of the primitives in the document. Furthermore, vector graphics are able to give extra structural information on how low-level elements group together to form high level shapes or structures. These merits of graphic vectors have not been fully leveraged in existing methods. To explore this data format, we target on the fundamental recognition tasks: object localization and classification. We propose an efficient CNN-free pipeline that does not render the graphic into pixels (i.e. rasterization), and takes textual document of the vector graphics as input, called YOLaT (You Only Look at Text). YOLaT builds multi-graphs to model the structural and spatial information in vector graphics, and a dual-stream graph neural network is proposed to detect objects from the graph. Our experiments show that by directly operating on vector graphics, YOLaT outperforms raster-graphic based object detection baselines in terms of both average precision and efficiency. Code is available at https://github.com/microsoft/YOLaT-VectorGraphicsRecognition

OPUS - University of Technology Sydney