Search CORE

583 research outputs found

Speaker Identification and Spoken word Recognition in Noisy Environment using Different Techniques

Author: Shaik Shafee, Prof. B. Anuradha
Publication venue: 'Auricle Technologies, Pvt., Ltd.'
Publication date: 30/06/2016
Field of study

In this work, an attempt is made to design ASR systems through software/computer programs which would perform Speaker Identification, Spoken word recognition and combination of both speaker identification and Spoken word recognition in general noisy environment. Automatic Speech Recognition system is designed for Limited vocabulary of Telugu language words/control commands. The experiments are conducted to find the better combination of feature extraction technique and classifier model that will perform well in general noisy environment (Home/Office environment where noise is around 15-35 dB). A recently proposed features extraction technique Gammatone frequency coefficients which is reported as the best fit to the human auditory system is chosen for the experiments along with the more common feature extraction techniques MFCC and PLP as part of Front end process (i.e. speech features extraction). Two different Artificial Neural Network classifiers Learning Vector Quantization (LVQ) neural networks and Radial Basis Function (RBF) neural networks along with Hidden Markov Models (HMMs) are chosen for the experiments as part of Back end process (i.e. training/modeling the ASRs). The performance of different ASR systems that are designed by utilizing the 9 different combinations (3 feature extraction techniques and 3 classifier models) are analyzed in terms of spoken word recognition and speaker identification accuracy success rate, design time of ASRs, and recognition / identification response time .The testing speech samples are recorded in general noisy conditions i.e.in the existence of air conditioning noise, fan noise, computer key board noise and far away cross talk noise. ASR systems designed and analyzed programmatically in MATLAB 2013(a) Environment

International Journal on Recent and Innovation Trends in Computing and Communication

Arabic digits speech recognition and speaker identification in noisy environment using a hybrid model of VQ and GMM

Author: Frikel Miloud
Ouisaadane Abdelkbir
Safi Said
Publication venue: 'Universitas Ahmad Dahlan'
Publication date: 01/08/2020
Field of study

This paper presents an automatic speaker identification and speech recognition for Arabic digits in noisy environment. In this work, the proposed system is able to identify the speaker after saving his voice in the database and adding noise. The mel frequency cepstral coefficients (MFCC) is the best approach used in building a program in the Matlab platform; also, the quantization is used for generating the codebooks. The Gaussian mixture modelling (GMM) algorithms are used to generate template, feature-matching purpose. In this paper, we have proposed a system based on MFCC-GMM and MFCC-VQ Approaches on the one hand and by using the Hybrid Approach MFCC-VQ-GMM on the other hand for speaker modeling. The White Gaussian noise is added to the clean speech at several signal-to-noise ratio (SNR) levels to test the system in a noisy environment. The proposed system gives good results in recognition rate

TELKOMNIKA (Telecommunication Computing Electronics and Control)

UAD Journal Management System

Cross-Word Arabic Pronunciation Variation Modeling Using Part of Speech Tagging

Author: AbuZeina Dia
Al-Muhtaseb Husni
Elshafei Moustafa
Publication venue: 'IntechOpen'
Publication date: 28/11/2012
Field of study

IntechOpen

Arabic Isolated Word Speaker Dependent Recognition System

Author: Elkourd Amer.m
Publication venue: الجامعة الإسلامية - غزة
Publication date: 01/01/2014
Field of study

In this thesis we designed a new Arabic isolated word speaker dependent recognition system based on a combination of several features extraction and classifications techniques. Where, the system combines the methods outputs using a voting rule. The system is implemented with a graphic user interface under Matlab using G62 Core I3/2.26 Ghz processor laptop. The dataset used in this system include 40 Arabic words recorded in a calm environment with 5 different speakers using laptop microphone. Each speaker will read each word 8 times. 5 of them are used in training and the remaining are used in the test phase. First in the preprocessing step we used an endpoint detection technique based on energy and zero crossing rates to identify the start and the end of each word and remove silences then we used a discrete wavelet transform to remove noise from signal. In order to accelerate the system and reduce the execution time we make the system first to recognize the speaker and load only the reference model of that user. We compared 5 different methods which are pairwise Euclidean distance with MelFrequency cepstral coefficients (MFCC), Dynamic Time Warping (DTW) with Formants features, Gaussian Mixture Model (GMM) with MFCC, MFCC+DTW and Itakura distance with Linear Predictive Coding features (LPC) and we got a recognition rate of 85.23%, 57% , 87%, 90%, 83% respectively. In order to improve the accuracy of the system, we tested several combinations of these 5 methods. We find that the best combination is MFCC | Euclidean + Formant | DTW + MFCC | DTW + LPC | Itakura with an accuracy of 94.39% but with large computation time of 2.9 seconds. In order to reduce the computation time of this hybrid, we compare several subcombination of it and find that the best performance in trade off computation time is by first combining MFCC | Euclidean + LPC | Itakura and only when the two methods do not match the system will add Formant | DTW + MFCC | DTW methods to the combination, where the average computation time is reduced to the half to 1.56 seconds and the system accuracy is improved to 94.56%. Finally, the proposed system is good and competitive compared with other previous researches

Institutional Repository of the Islamic University of Gaza

Recent Improvements in the BBN OCR System

Author: Issam Bazzi
John Makhoul
Kornai András
Prem Natarajan
Richard Schwartz
Zhidong Lu
Publication venue
Publication date: 01/01/1999
Field of study

SZTAKI Publication Repository

A prior case study of natural language processing on different domain

Author: J. Shruthi
Swamy Suma
Publication venue: 'Institute of Advanced Engineering and Science'
Publication date: 01/10/2020
Field of study

In the present state of digital world, computer machine do not understand the human’s ordinary language. This is the great barrier between humans and digital systems. Hence, researchers found an advanced technology that provides information to the users from the digital machine. However, natural language processing (i.e. NLP) is a branch of AI that has significant implication on the ways that computer machine and humans can interact. NLP has become an essential technology in bridging the communication gap between humans and digital data. Thus, this study provides the necessity of the NLP in the current computing world along with different approaches and their applications. It also, highlights the key challenges in the development of new NLP model

ZENODO

Institute of Advanced Engineering and Science

Automatic Identity Recognition Using Speech Biometric

Author: Abushariah Mohammad A. M.
Alqudah Assal A. M.
Publication venue: 'European Scientific Institute, ESI'
Publication date: 28/04/2016
Field of study

Biometric technology refers to the automatic identification of a person using physical or behavioral traits associated with him/her. This technology can be an excellent candidate for developing intelligent systems such as speaker identification, facial recognition, signature verification...etc. Biometric technology can be used to design and develop automatic identity recognition systems, which are highly demanded and can be used in banking systems, employee identification, immigration, e-commerce…etc. The first phase of this research emphasizes on the development of automatic identity recognizer using speech biometric technology based on Artificial Intelligence (AI) techniques provided in MATLAB. For our phase one, speech data is collected from 20 (10 male and 10 female) participants in order to develop the recognizer. The speech data include utterances recorded for the English language digits (0 to 9), where each participant recorded each digit 3 times, which resulted in a total of 600 utterances for all participants. For our phase two, speech data is collected from 100 (50 male and 50 female) participants in order to develop the recognizer. The speech data is divided into text-dependent and text-independent data, whereby each participant selected his/her full name and recorded it 30 times, which makes up the text-independent data. On the other hand, the text-dependent data is represented by a short Arabic language story that contains 16 sentences, whereby every sentence was recorded by every participant 5 times. As a result, this new corpus contains 3000 (30 utterances * 100 speakers) sound files that represent the text-independent data using their full names and 8000 (16 sentences * 5 utterances * 100 speakers) sound files that represent the text-dependent data using the short story. For the purpose of our phase one of developing the automatic identity recognizer using speech, the 600 utterances have undergone the feature extraction and feature classification phases. The speech-based automatic identity recognition system is based on the most dominating feature extraction technique, which is known as the Mel-Frequency Cepstral Coefficient (MFCC). For feature classification phase, the system is based on the Vector Quantization (VQ) algorithm. Based on our experimental results, the highest accuracy achieved is 76%. The experimental results have shown acceptable performance, but can be improved further in our phase two using larger speech data size and better performance classification techniques such as the Hidden Markov Model (HMM)

European Scientific Journal, ESJ

European Scientific Journal (European Scientific Institute)

Continuous Density Hidden Markov Model for Hindi Speech Recognition

Author: . Aruna Jain3
. S S Agrawal
. Shweta Sinha
Publication venue: GSTF Journal on Computing (JoC)
Publication date: 26/08/2014
Field of study

State of the art automatic speech recognitionsystem uses Mel frequency cepstral coefficients as featureextractor along with Gaussian mixture model for acousticmodeling but there is no standard value to assign number ofmixture component in speech recognition process.Currentchoice of mixture component is arbitrary with littlejustification. Also the standard set for European languagescan not be used in Hindi speech recognition due to mismatchin database size of the languages.Parameter estimation withtoo many or few component may inappropriately estimatethe mixture model. Therefore, number of mixture isimportant for initial estimation of expectation maximizationprocess. In this research work, the authors estimate numberof Gaussian mixture component for Hindi database basedupon the size of vocabulary.Mel frequency cepstral featureand perceptual linear predictive feature along with itsextended variations with delta-delta-delta feature have beenused to evaluate this number based on optimal recognitionscore of the system . Comparitive analysis of recognitionperformance for both the feature extraction methods onmedium size Hindi database is also presented in thispaper.HLDA has been used as feature reduction techniqueand also its impact on the recognition score has beenhighlighted

GSTF Digital Library (GSTF-DL): Open Journal Systems (Global Science and Technology Forum)