Search CORE

601 research outputs found

BaDLAD: A Large Multi-Domain Bengali Document Layout Analysis Dataset

Author: Ahmed Intesur
Ansary Md. Nazmuddoha
Chowdhury Sayma Sultana
Dhruvo Shahriar Elahi
Dip Souhardya Saha
Emon Mahfuzur Rahman
Haque Md. Rezwanul
Hasan Md. Rakibul
Hossen Syed Mobassir
Humayun Ahmed Imtiaz
Meghla Marsia Haque
Pavel Akib Hasan
Rakib Fazle Rabbi
Reasat Tahsin
Sadeque Farig
Shihab Md. Istiak Hossain
Sushmit Asif Shahriyar
Publication venue
Publication date: 10/03/2023
Field of study

While strides have been made in deep learning based Bengali Optical Character Recognition (OCR) in the past decade, the absence of large Document Layout Analysis (DLA) datasets has hindered the application of OCR in document transcription, e.g., transcribing historical documents and newspapers. Moreover, rule-based DLA systems that are currently being employed in practice are not robust to domain variations and out-of-distribution layouts. To this end, we present the first multidomain large Bengali Document Layout Analysis Dataset: BaDLAD. This dataset contains 33,695 human annotated document samples from six domains - i) books and magazines, ii) public domain govt. documents, iii) liberation war documents, iv) newspapers, v) historical newspapers, and vi) property deeds, with 710K polygon annotations for four unit types: text-box, paragraph, image, and table. Through preliminary experiments benchmarking the performance of existing state-of-the-art deep learning architectures for English DLA, we demonstrate the efficacy of our dataset in training deep learning based Bengali document digitization models

arXiv.org e-Print Archive

A Vietnamese Handwritten Text Recognition Pipeline for Tetanus Medical Records

Author: Bui Triet
Dinh Minh N
Le Mau Toan
Mai Minh
Nguyen Nhan
Tran Long
Vo Tan Hoang
Publication venue: AIS Electronic Library (AISeL)
Publication date: 11/12/2023
Field of study

Machine learning techniques are successful for optical character recognition tasks, especially in recognizing handwriting. However, recognizing Vietnamese handwriting is challenging with the presence of extra six distinctive tonal symbols and vowels. Such a challenge is amplified given the handwriting of health workers in an emergency care setting, where staff is under constant pressure to record the well-being of patients. In this study, we aim to digitize the handwriting of Vietnamese health workers. We develop a complete handwritten text recognition pipeline that receives scanned documents, detects, and enhances the handwriting text areas of interest, transcribes the images into computer text, and finally auto-corrects invalid words and terms to achieve high accuracy. From experiments with medical documents written by 30 doctors and nurses from the Tetanus Emergency Care unit at the Hospital for Tropical Diseases, we obtain promising results of 2% and 12% for Character Error Rate and Word Error Rate, respectively

AIS Electronic Library (AISeL)

DAN: a Segmentation-free Document Attention Network for Handwritten Document Recognition

Author: Chatelain Clément
Coquenet Denis
Paquet Thierry
Publication venue
Publication date: 01/08/2022
Field of study

Unconstrained handwritten text recognition is a challenging computer vision task. It is traditionally handled by a two-step approach, combining line segmentation followed by text line recognition. For the first time, we propose an end-to-end segmentation-free architecture for the task of handwritten document recognition: the Document Attention Network. In addition to text recognition, the model is trained to label text parts using begin and end tags in an XML-like fashion. This model is made up of an FCN encoder for feature extraction and a stack of transformer decoder layers for a recurrent token-by-token prediction process. It takes whole text documents as input and sequentially outputs characters, as well as logical layout tokens. Contrary to the existing segmentation-based approaches, the model is trained without using any segmentation label. We achieve competitive results on the READ 2016 dataset at page level, as well as double-page level with a CER of 3.43% and 3.70%, respectively. We also provide results for the RIMES 2009 dataset at page level, reaching 4.54% of CER. We provide all source code and pre-trained model weights at https://github.com/FactoDeepLearning/DAN

arXiv.org e-Print Archive

HAL - Normandie Université

WordSup: Exploiting Word Annotations for Character based Text Detection

Author: Ding Errui
Han Junyu
Hu Han
Luo Yuxuan
Wang Yuzhuo
Zhang Chengquan
Publication venue
Publication date: 22/08/2017
Field of study

Imagery texts are usually organized as a hierarchy of several visual elements, i.e. characters, words, text lines and text blocks. Among these elements, character is the most basic one for various languages such as Western, Chinese, Japanese, mathematical expression and etc. It is natural and convenient to construct a common text detection engine based on character detectors. However, training character detectors requires a vast of location annotated characters, which are expensive to obtain. Actually, the existing real text datasets are mostly annotated in word or line level. To remedy this dilemma, we propose a weakly supervised framework that can utilize word annotations, either in tight quadrangles or the more loose bounding boxes, for character detector training. When applied in scene text detection, we are thus able to train a robust character detector by exploiting word annotations in the rich large-scale real scene text datasets, e.g. ICDAR15 and COCO-text. The character detector acts as a key role in the pipeline of our text detection engine. It achieves the state-of-the-art performance on several challenging scene text detection benchmarks. We also demonstrate the flexibility of our pipeline by various scenarios, including deformed text detection and math expression recognition.Comment: 2017 International Conference on Computer Visio

arXiv.org e-Print Archive

Crossref

Combining Unsupervised, Supervised, and Rule-based Algorithms for Text Mining of Electronic Health Records - A Clinical Decision Support System for Identifying and Classifying Allergies of Concern for Anesthesia During Surgery

Author: Berge Geir Thore
Granmo Ole-Christoffer
Tveit Tor Oddbjørn
Publication venue: AIS Electronic Library (AISeL)
Publication date: 27/09/2017
Field of study

Undisclosed allergic reactions of patients are a major risk when undertaking surgeries in hospitals. We present our early experience and preliminary findings for a Clinical Decision Support System (CDSS) being developed in a Norwegian Hospital Trust. The system incorporates unsupervised and supervised machine learning algorithms in combination with rule-based algorithms to identify and classify allergies of concern for anesthesia during surgery. Our approach is novel in that it utilizes unsupervised machine learning to analyze large corpora of narratives to automatically build a clinical language model containing words and phrases of which meanings and relative meanings are also learnt. It further implements a semi-automatic annotation scheme for efficient and interactive machine-learning, which to a large extent eliminates the substantial manual annotation (of clinical narratives) effort necessary for the training of supervised algorithms. Validation of system performance was performed through comparing allergies identified by the CDSS with a manual reference standard

AIS Electronic Library (AISeL)

XDOCS: An Application to Index Historical Documents

Author: BOLELLI FEDERICO
BORGHI GUIDO
GRANA Costantino
Publication venue: 'Springer Science and Business Media LLC'
Publication date: 01/01/2018
Field of study

Dematerialization and digitalization of historical documents are key elements for their availability, preservation and diffusion. Unfortunately, the conversion from handwritten to digitalized documents presents several technical challenges. The XDOCS project is created with the main goal of making available and extending the usability of historical documents for a great variety of audience, like scholars, institutions and libraries. In this paper the core elements of XDOCS, i.e. page dewarping and word spotting technique, are described and two new applications, i.e. annotation/indexing and search tool, are presented

Archivio istituzionale della ricerca - Alma Mater Studiorum Università di Bologna

Archivio istituzionale della ricerca - Università di Modena e Reggio Emilia

ArchMine: Learning from non-machine-readable documents for additional insights

Author: Mariana Ferreira Dias
Publication venue
Publication date: 17/03/2023
Field of study

Repositório Aberto da Universidade do Porto